Google Soli proved that a millimetre-wave radar small enough for consumer electronics could recognize a compact vocabulary of deliberate finger and hand motions by learning their changing radio-frequency signatures—without reconstructing a camera-like image of the hand. In the 2016 study, its filtered classifier achieved 92.10% per-gesture accuracy on 1,000 test gestures from five participants. A separate robotic-target experiment measured displacement with 0.43 mm average RMS error. Those are important results, but they do not establish universal, user-independent gesture recognition in uncontrolled daily life.
That distinction is the heart of the Soli story. The breakthrough was not that radar suddenly produced a detailed picture of individual fingers. It was that a compact 60 GHz sensor could capture motion quickly and precisely enough for software to recognize useful temporal patterns even when its conventional spatial resolution was comparatively coarse.
Short answer: Soli demonstrated a new interaction architecture: illuminate the whole hand, measure the composite reflection at very high update rates, convert its changes into motion features, and classify a small set of intentional gestures. It proved technical feasibility—not a solved replacement for touchscreens.
The paper at a glance
| Paper | Soli: Ubiquitous Gesture Sensing with Millimeter Wave Radar |
|---|---|
| Authors | Jaime Lien, Nicholas Gillian, M. Emre Karagozler, Patrick Amihood, Carsten Schwesig, Erik Olson, Hakim Raja and Ivan Poupyrev |
| Published | ACM Transactions on Graphics, Volume 35, Issue 4, Article 142, July 2016 |
| Core question | Can radar be redesigned as a miniature, low-power input sensor for fine hand gestures? |
| Evaluated gestures | Virtual Button, Virtual Slider, horizontal swipe and vertical swipe |
| Main reported result | 92.10% filtered per-gesture accuracy on 1,000 second-session test gestures |
| Correct interpretation | A strong controlled proof of concept with five known participants—not evidence of open-world recognition |
The original paper and its full methods are available through the ACM DOI record.
Why gesture sensing needed a different kind of radar
Traditional radar systems were built to detect comparatively large, often rigid targets such as aircraft, ships and vehicles. Fine hand interaction reverses many of those assumptions. The target is close, deformable and made of multiple moving parts. The motion may be only a fraction of a millimetre, while the sensor must fit inside a phone, watch or appliance and operate within tight power and computation limits.
Cameras and depth sensors can estimate hand shape, but they depend on an optical view and can require substantial spatial processing. Capacitive sensing is compact, but its useful free-space range is limited. Soli’s researchers asked whether radar could occupy a different point in the design space: compact, fast, insensitive to lighting and capable of measuring subtle motion without capturing an ordinary image.
The team progressed from desktop-scale 57–64 GHz FMCW hardware and 3–10 GHz impulse-radar prototypes to integrated 60 GHz devices. The 2016 paper described a 12 × 12 mm SiGe FMCW chip and a 9 × 9 mm CMOS direct-sequence spread-spectrum chip, both with antennas integrated into the package or die. It also described two transmit and four receive elements, a broad approximately 150-degree beam and prototype chip power consumption of about 300 mW with adaptive duty cycling. These were research-platform figures, not universal specifications for every later Soli product (Lien et al., 2016).
The key idea: recognize motion without imaging the hand
A conventional imaging instinct would be to form a narrow beam, separate the fingers spatially and reconstruct a hand skeleton. Soli deliberately avoided making that the primary task. Its wide beam illuminated the whole hand. Reflections from fingertips, joints and other scattering regions arrived as a complex superposition rather than a neat picture.
The system then looked at how that composite return changed over time. A thumb tapping an index finger, a thumb sliding along a finger and a hand sweeping through the sensing region generate different combinations of range, radial velocity, phase and reflected energy. Machine learning can discriminate those signatures even when the sensor cannot spatially resolve every finger.
This is analogous to recognizing a musical instrument from its sound rather than from a photograph. The waveform does not provide a literal picture of the instrument, but its changing structure can still be distinctive enough to classify.

Range resolution is not displacement sensitivity
This is the most frequently missed technical point in descriptions of Soli. For an ideal radar, range resolution is approximately ΔR = c/(2B), where c is wave propagation speed and B is transmitted bandwidth. At 7 GHz of bandwidth, that gives about 2.14 cm. Two fingertips separated only along the range direction by a few millimetres would therefore not appear as two conventionally resolved range targets.
Displacement sensitivity asks a different question: how precisely can the system detect a change in the position of one coherent return? A radial displacement Δx changes the round-trip phase by approximately Δφ = 4πΔx/λ. At 60 GHz, the wavelength is about 5 mm, so a 1 mm radial displacement produces roughly 0.8π radians, or 144 degrees, of phase change.
The Soli paper estimated an ideal displacement limit of about 0.3 mm from its phase jitter. In a validation experiment, a robotic arm moved a metal plate at twelve velocities from 20 to 200 mm/s while a laser range finder provided a reference. The radar’s average RMS displacement error was 0.43 mm. That supports sub-millimetre displacement tracking for the tested target and setup; it is not the same as resolving two objects 0.43 mm apart, nor is it a fingertip-position error measured during the gesture study (Lien et al., 2016).
For a broader treatment of this distinction, see ArthaVedya’s guide to choosing between 24, 60 and 77 GHz radar for human sensing.
How the Soli processing pipeline worked
The paper’s contribution was larger than a single classifier. It described an end-to-end stack connecting the radar waveform, integrated hardware, signal transformations, features, gesture vocabulary and application interface.
- Transmit and receive: a broad beam illuminates the hand, and the receiver records the combined reflections from multiple moving scattering regions.
- Sample rapidly: the proposed design used radar repetition frequencies between 1 and 10 kHz. High temporal sampling captures phase and velocity changes produced by fine motion.
- Transform the signal: the software forms representations including complex I/Q data, range profiles, Doppler profiles, range–Doppler maps and spectrogram-like micro-Doppler views.
- Extract compact features: examples include range, velocity, acceleration, velocity centroid, total energy, moving energy and fine displacement.
- Classify over time: a random-forest model estimates gesture likelihoods, while temporal and contextual filtering stabilizes the output and reduces isolated false detections.
The hardware-abstraction layer was strategically important. Soli transformed hardware-specific measurements into common representations so higher-level applications would not have to be rewritten around each radar architecture. The team demonstrated the approach across FMCW, phase-modulated spread-spectrum and impulse-radar prototypes.
Micro-Doppler is especially useful when different parts of a hand move at different radial velocities. It describes motion-induced frequency structure over time; it is not a video silhouette. Readers working with this representation can continue with the practical guide to choosing STFT parameters for radar micro-Doppler.
For a deeper physical interpretation of how multiple moving body parts create these signatures, see Micro-Doppler Explained: How Radar Sees Arms, Legs, and Human Motion.
Why “Virtual Tools” mattered as much as the sensor
Soli was not designed around an unlimited dictionary of arbitrary hand signs. The researchers instead proposed “Virtual Tools”: small gestures that imitate familiar physical controls. A thumb and index finger could act like a button, the thumb could move along the side of the index finger like a slider, or two fingers could manipulate an imagined dial.
This was an interaction-design decision as well as an engineering one. Small finger motions exploit the sensor’s temporal sensitivity, reduce the fatigue and social awkwardness of large arm gestures, and give the user a natural reference through proprioception. When two fingers touch or move against each other, the hand also supplies its own tactile cue even though the device provides no physical control surface.
The evaluated vocabulary was narrower than the full design vision. The experiment tested four classes: Virtual Button, Virtual Slider, horizontal swipe and vertical swipe. The Virtual Dial was discussed as an interaction concept but was not one of the four classes reported in the main recognition evaluation.
What the experiment actually tested
Five participants sat in front of a stationary sensor and performed each of the four gestures 50 times at different positions within 30 cm. They completed two recording sessions, with a break between them. The first session from every participant formed the training set; the second session from those same participants formed the test set. Background movements and transitions into and out of valid gestures were also recorded (Lien et al., 2016).
Across both sessions, the dataset contained 2,000 gesture examples: 1,000 for training and 1,000 for testing. The researchers sampled 2,500 raw radar frames per second on two virtual channels, computed transformations every ten frames, and classified features accumulated over a ten-transformation window of approximately 40 ms.
The random forest used 50 trees with a maximum depth of 10. Its raw decisions were then passed through a Bayesian temporal filter. This matters because the best headline result belongs to the complete filtered system, not to the random forest alone.
| Metric | Raw classifier | After temporal filtering |
|---|---|---|
| Per-sample accuracy | 73.64% | 78.22% |
| Per-gesture accuracy | 86.90% | 92.10% |
Per-sample accuracy asks whether each time sample received the correct label. Per-gesture accuracy asks whether an entire gesture event was recognized correctly. The lower sample-level score partly reflects uncertainty near the beginning and end of a gesture, where frames can resemble background movement. Temporal filtering improved both metrics by using continuity rather than treating every instant as independent.
The train–test split was better than randomly mixing highly correlated frames from the same recording, a practice the authors said produced unrealistically high results. However, it was not a participant-independent test: all five people contributed to both training and testing. Modern human-sensing studies should report subject-independent evaluation when they claim generalization to unseen users. ArthaVedya’s radar-based human activity recognition guide explains why this split choice can change the scientific conclusion.
What “more than 10,000 frames per second” meant
The paper’s abstract states that Soli ran at more than 10,000 frames per second on embedded hardware. That sentence is easy to misread. In an unthrottled software benchmark, the optimized fine-displacement pipeline processed 11,500 frames/s on a Raspberry Pi 2 and 18,000 frames/s on a Snapdragon 400. The complete gesture-recognition pipeline reached 1,480 and 2,880 frames/s on those platforms, respectively (Lien et al., 2016).
Those figures measured processing throughput using recorded input and one CPU thread at full utilization. They were not the rate at which a person received 10,000 independent gesture decisions. In the recognition experiment, raw acquisition was 2,500 frames/s and transformations were produced at 250 per second. The benchmark nevertheless supported an important engineering claim: the pipeline was light enough for real-time embedded execution with substantial timing headroom.
What Soli genuinely proved
| Demonstrated claim | Evidence in the paper | Boundary |
|---|---|---|
| Millimetre-wave radar can sense fine motion at close range | 0.43 mm average RMS displacement error against a laser reference | Measured with a moving metal plate, not free-moving fingertips |
| A compact radar can support gesture input | Integrated 9 × 9 mm and 12 × 12 mm 60 GHz radar chips and embedded processing | Research prototypes; product constraints still matter |
| Temporal signatures can substitute for detailed hand imaging | Four gesture classes recognized from I/Q, range–Doppler and motion features | No full hand skeleton or arbitrary pose reconstruction |
| A small vocabulary can work in a continuous stream | 92.10% filtered per-gesture accuracy with background motion included | Five instructed participants in a controlled setup |
| The processing can run on embedded hardware | Real-time throughput on Raspberry Pi 2 and Snapdragon 400 platforms | Benchmark throughput is not end-to-end interaction accuracy |
Together, these results justified treating radar as an interaction sensor rather than merely shrinking a conventional tracking radar. Soli’s real conceptual advance was co-design: the hardware, temporal sampling, signal representations, feature set and gesture vocabulary were chosen together.
What the paper did not prove
- It did not prove user-independent recognition. The same five participants appeared in the training and test sessions.
- It did not prove a large or open-ended gesture language. The main evaluation used four intentional gestures, and the paper itself noted that reliability generally decreases as a gesture set grows.
- It did not prove robustness in every environment. The authors explicitly left clutter, coupling, multipath, fading, interference and occlusion as open research problems.
- It did not reconstruct detailed hand anatomy. Classification relied on learned motion signatures rather than identifying every finger or recovering a complete skeletal pose.
- It did not establish that radar should replace touch. False positives, discoverability, fatigue, feedback, power, radio regulation and application context still determine whether touchless input is useful.
These limitations do not diminish the paper. They define what kind of contribution it was: a persuasive systems proof of concept and a new design framework, not a final benchmark for unconstrained deployment.
What happened when Soli left the laboratory
Productization exposed constraints that a classifier table cannot capture. In December 2018, the US Federal Communications Commission granted Google a waiver for Soli operation in the 57–64 GHz band at higher power than the general limit then applicable to that class of sensor, subject to specified power and duty-cycle conditions. The FCC order records Google’s argument that the lower limit led to missed motions and fewer effective interactions. This is a valuable engineering lesson: sensing performance depends on the permitted waveform and link budget as well as the learning algorithm.
Google then placed Soli in the Pixel 4. Its official 2019 announcement described a deliberately restricted set of Motion Sense functions: skipping songs, snoozing alarms, silencing calls and detecting that a user was reaching for the phone so face-unlock hardware could activate. Google also stated that the Soli sensor data was processed on the phone and was not saved or shared with its other services. These details are documented in the Pixel 4 Motion Sense announcement.
The narrow product vocabulary is not evidence that the research failed. It is evidence that a reliable consumer interface must balance recognition, accidental activation, learnability, power and regional authorization. The 2016 paper proposed small “Virtual Toolkits” for essentially the same reason: a compact, context-aware set can be more useful than an impressive but fragile catalogue of gestures.
Soli later illustrated a broader point in the second-generation Nest Hub: the same basic ability to detect subtle contactless motion could support bedside sleep sensing rather than explicit hand commands. A 2021 Google validation reported 87% epoch-by-epoch sleep–wake accuracy in healthy sleepers, with 96% of sleep epochs and 55% of wake epochs detected correctly. The asymmetry is as informative as the overall number. It shows both the value of radar micro-motion and the need to inspect class-specific performance. The study summary and paper are available from Google Research.
Lessons for radar and machine-learning researchers
1. Design the vocabulary around the sensor
A model cannot recover distinctions that the measurement does not preserve reliably. Gestures should produce separable radial-motion, phase, energy or timing patterns and should also be easy for people to repeat. Soli’s Virtual Tools linked sensing physics to motor behaviour instead of treating gesture selection as an afterthought.
2. Treat temporal context as part of the system
The improvement from 86.90% raw to 92.10% filtered per-gesture accuracy was not free. It came from using neighbouring predictions and application context. Any comparison with another classifier should therefore specify whether smoothing, spotting logic and rejection rules are included.
3. Separate correlated samples before reporting accuracy
Thousands of frames from one gesture are not thousands of independent trials. Soli’s authors explicitly rejected random frame-level validation because correlation produced unrealistically high performance. For a modern study, splitting by participant, session, environment and device is often necessary to support claims of generalization.
4. Name every metric precisely
Range resolution, range accuracy, displacement precision, acquisition rate, transformation rate, processing throughput and event-level accuracy describe different properties. Combining them into a single claim such as “sub-millimetre radar at 10,000 fps” may sound impressive while obscuring what was actually measured.
5. Product constraints are part of sensing science
A laboratory model may be accurate yet fail as an interface because of false activations, limited feedback, RF coexistence, power consumption or an awkward gesture vocabulary. The Soli story is valuable precisely because it connects a paper architecture to regulatory review and real consumer devices.
The verdict
Soli: Ubiquitous Gesture Sensing with Millimeter Wave Radar deserves its landmark status, but for a more precise reason than the usual “radar can see your fingers” summary. It showed that fine interaction does not always require fine spatial imaging. A sensor can trade detailed geometry for high-rate, phase-sensitive temporal measurements and still recover a useful command.
The reported 92.10% gesture accuracy was credible within a small, controlled same-participant experiment. The 0.43 mm displacement result validated a physical sensing capability under a different metallic-target experiment. Embedded benchmarks showed that the processing architecture was practical. Later devices established that the sensing concept could leave the laboratory, although the deployed gesture set remained intentionally modest.
What Google Soli actually proved was not universal gesture understanding. It proved that a miniature radar could become a new kind of input device—and that temporal signal design, interaction design and machine learning had to be engineered as one system.
Key takeaways
- Soli recognized motion signatures; it did not need to form a camera-like hand image.
- Its approximately 2 cm nominal range resolution and 0.43 mm measured displacement error are different quantities.
- The best reported gesture result was 92.10% after temporal filtering on four gestures from five known participants.
- The same participants appeared in training and test sessions, so the study did not establish unseen-user generalization.
- The “more than 10,000 fps” result referred to embedded processing throughput for the fine-displacement pipeline, not 10,000 independent gesture decisions per second.
- Soli’s lasting lesson is co-design: sensing physics, processing, evaluation and human interaction must fit one another.
References
- Lien, J., Gillian, N., Karagozler, M. E., Amihood, P., Schwesig, C., Olson, E., Raja, H., & Poupyrev, I. (2016). Soli: Ubiquitous Gesture Sensing with Millimeter Wave Radar. ACM Transactions on Graphics, 35(4), Article 142.
- Federal Communications Commission. (2018). In the Matter of Google LLC: Request for Waiver of Section 15.255(c)(3), DA 18-1308.
- Barbello, B. (2019). (Don’t) hold the phone: New features coming to Pixel 4. Google.
- Dixon, M., Schneider, L., Yu, J., et al. (2021). Sleep-wake Detection With a Contactless, Bedside Radar Sleep Sensing System. Google.
