Egocentric video, gaze, inertial data, audio and peripheral physiology from one wearable system, timestamped against a single clock and built to work in the field.
Desmo looks and wears like a regular pair of glasses, with no cap, gel, or complex rig to strap on. Put them on, carry on with your day, and get hours of rich, labelled, multi-modal data.
On the frame: eye cameras, a forward scene camera, a four-mic array and head motion. A wrist band adds cardiac and thermal channels, and a portable puck records and runs inference on-device.
Capabilities
Hover to inspect · click to frame
Fig. 02 · Head unit, 3DHover lights a subsystem · click to focus · drag to rotate
Detail
Select a capability to see its specification.
Eye tracking
2 × mono NIR global shutter · 120 fps · 160° DFOV
Gaze, fixations, saccades to 500°/s, blinks and pupil diameter. Global shutter freezes the pupil mid-saccade, and 160° DFOV tolerates frame slippage. Eye-safe 850 nm illumination gives a corneal glint reference; extraction happens on the frame, so coordinates leave and raw eye video does not.
Forward scene video stamped against gaze, so fixations land on pixels. Global shutter removes rolling-shutter wobble during head turns, and 152° DFOV covers the full gaze range. Feeds object detection, OCR and saliency analysis.
Head motion
6-axis IMU · 200 Hz · bridge-mounted
Head orientation for gaze-in-world reconstruction, video stabilisation, and honest motion-artifact flags on the epochs that need them. Centred on the rigid bridge so temple flex never reads as motion.
Spatial audio
4-mic array · 68 dB SNR · 133 dB AOP
Beamformed speech, ambient audio and prosody for voice-stress analysis without clipping. Open-ear speakers deliver auditory stimuli in the field, hardware-capped at 85 dB and timestamped so they can be regressed out.
On-device processing
1.0 TOPS NPU · H.264 / H.265
Pupillometry and video encoding run on the glasses. What crosses the tether is compressed video and extracted signals, not raw imagery.
Touch input
Capacitive slider
Tap, swipe and scrub on the temple for event marking and control, so there is no phone in the participant’s hands mid-task.
Power and connectivity
3–4 h · 1000 mAh backup · 64 GB
Powered from the puck over one USB-C tether, with Wi-Fi 6 and BLE 5.2 for control. If the tether disconnects, the on-board cell and local storage buffer the recording so nothing is lost mid-session.
PeripheralsSame session, same clock
Wristband
Sensing
Peripheral autonomic channels: beat-to-beat cardiac intervals, skin temperature and wrist motion. Each sample is timestamped against a ±1 ppm RTC at source before transmission over BLE.
Heart activity
Multi-wavelength PPG
Skin temperature
±0.1 °C @ 1 Hz
Motion
6-axis IMU @ 200 Hz
Puck
Compute & Storage
Portable recorder and inference host. Its system clock is the master every stream is stamped against, with 500 GB of local storage for raw capture, Wi-Fi 6E and optional LTE.
Clock
Single master, <10 ms
Compute
16 GB LPDDR5
Runtime
3–4 h continuous
02 / Synchronization
Synced to follow a single clock
Most labs run eye tracking, scene video and physiology on separate systems and align them afterwards. Desmo records every channel against one clock on the puck.
A single stimulus is answered across three orders of magnitude: an orienting saccade within a couple of hundred milliseconds, vascular response over minutes. The range is what makes one instrument hard to build. The fast end is what makes the clock hard: the saccade and the pupil dilation that follows it sit roughly 100 ms apart, and resolving that ordering sets the requirement. Everything slower comes free.
Response
10 ms100 ms1 s10 s100 s
Saccade to stimulus
Eye camera · 100–200 ms
Pupil dilation (TEPR)
Pupillometry · 200–400 ms
Cardiac orienting response
PPG · 1–3 s
Vagal withdrawal
HRV, HF band · 5–10 s
Peripheral vasoconstriction
Skin temp · 15 s–2 min
Head unitWrist unitLog time axis · head unit → puck <1 ms (USB) · band → puck 1–5 ms (BLE, stamped at source)
03 / Robotics data
From capture to training-ready
A single wearer produces synchronised egocentric video, 200 Hz head and wrist inertial data, and gaze. Those are the observation, action and intent channels a policy needs from a human demonstration.
The constraint in robot learning is no longer video volume. It is the fraction of captured hours that pass a lab's quality bar, and the lag before they are labelled. Desmo is built to raise the first and remove the second.
Curation
Quality gate at capture time
Gaze stability and pupil response score every frame as it is recorded. Footage that would fail a hand-pose quality bar is flagged on the puck before it ships, not after annotation.
GazePupillometryOn-device compute
Labels
Measured, not estimated
Egocentric video at 30 fps underdetermines limb trajectory. Wrist trajectory comes from 200 Hz IMU instead, attended target from gaze, head pose from visual-inertial tracking. Every label is read off the sensors on one clock, with no video-estimation pipeline and no post-hoc alignment.
Wrist IMUHead IMUGaze
Annotation
Attention, annotated for free
The eye fixates a target several hundred milliseconds before the hand moves. Every session ships with gaze-grounded object labels: which object, from when, to when, with no annotation cost and no annotation lag.
GazeScene cam
Spatial
Pose without the alignment tax
Visual-inertial odometry needs IMU and video on a common time base. Both leave the recorder already stamped, so head pose and scene reconstruction need no temporal calibration step.
Head IMUScene cam
Delivery
Your schema, our export
Sessions are delivered as VRS, LeRobot, or RLDS, with per-file health checks. PII is blurred on-device before anything leaves the puck.
VRSLeRobotRLDSPII blur
Adjacent fieldsSame recording, same clock
Human factors
Workload, fatigue and attention lapses quantified during the task itself, in the cab, the cockpit or on the floor.
Teleoperation
Gaze and head pose as a control channel, logged alongside the scene the operator was working from.
Assistive AR
On-device context from scene video and gaze, with the compute to run inference at capture time rather than after the session.
04 / Research
Our claims come with citations.
Read more about our studies beyond the platform.
arXiv:2603.04098cs.CV · Mar 2026
Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning
Gaze confidence indexes visual stability; pupil response indexes information novelty. Gating on the first and ranking on the second retains 10% of frames at full-stream activity-recognition accuracy, with no model inference at capture time.
Riding Brainwaves in LLM Space: Understanding Activation Patterns Using Individual Neural Signatures
Frozen LLM representations contain person-specific neural directions. Per-participant linear probes on word-level EEG outperform a population probe ninefold on high-gamma power, and the directions do not transfer between individuals.
We’re looking for a small group of research teams to design and validate Desmo with us in real-world studies. Tell us about your workflow, sensing needs, and what you want to learn—we’ll explore a partnership together.
Partnership request received
Thanks—we’ve received your note. The Desmo team will follow up to explore fit and next steps.