Design partner program open

See the world, inside and out.

Egocentric video, gaze, inertial data, audio and peripheral physiology from one wearable system, timestamped against a single clock and built to work in the field.

Become a design partner Learn More
Illustration of hands assembling electronics with tracked hand landmarks
Fig. 01 · Desmo system Head unit + peripherals
8
Concurrent signal streams
<10 ms
Cross-device alignment, worst case
120 fps
Global-shutter eye tracking, 160° DFOV
01 / System

Several data streams. One wearable system.

Desmo looks and wears like a regular pair of glasses, with no cap, gel, or complex rig to strap on. Put them on, carry on with your day, and get hours of rich, labelled, multi-modal data.

On the frame: eye cameras, a forward scene camera, a four-mic array and head motion. A wrist band adds cardiac and thermal channels, and a portable puck records and runs inference on-device.

Capabilities
Hover to inspect · click to frame
Fig. 02 · Head unit, 3D Hover lights a subsystem · click to focus · drag to rotate
Detail
Select a capability to see its specification.
Peripherals Same session, same clock
Desmo wristband

Wristband

Sensing

Peripheral autonomic channels: beat-to-beat cardiac intervals, skin temperature and wrist motion. Each sample is timestamped against a ±1 ppm RTC at source before transmission over BLE.

Heart activity
Multi-wavelength PPG
Skin temperature
±0.1 °C @ 1 Hz
Motion
6-axis IMU @ 200 Hz
Desmo portable compute puck

Puck

Compute & Storage

Portable recorder and inference host. Its system clock is the master every stream is stamped against, with 500 GB of local storage for raw capture, Wi-Fi 6E and optional LTE.

Clock
Single master, <10 ms
Compute
16 GB LPDDR5
Runtime
3–4 h continuous
02 / Synchronization

Synced to follow a single clock

Most labs run eye tracking, scene video and physiology on separate systems and align them afterwards. Desmo records every channel against one clock on the puck.

A single stimulus is answered across three orders of magnitude: an orienting saccade within a couple of hundred milliseconds, vascular response over minutes. The range is what makes one instrument hard to build. The fast end is what makes the clock hard: the saccade and the pupil dilation that follows it sit roughly 100 ms apart, and resolving that ordering sets the requirement. Everything slower comes free.

Response
10 ms 100 ms 1 s 10 s 100 s
Saccade to stimulus
Eye camera · 100–200 ms
Pupil dilation (TEPR)
Pupillometry · 200–400 ms
Cardiac orienting response
PPG · 1–3 s
Vagal withdrawal
HRV, HF band · 5–10 s
Peripheral vasoconstriction
Skin temp · 15 s–2 min
Head unit Wrist unit Log time axis · head unit → puck <1 ms (USB) · band → puck 1–5 ms (BLE, stamped at source)
03 / Robotics data

From capture to training-ready

A single wearer produces synchronised egocentric video, 200 Hz head and wrist inertial data, and gaze. Those are the observation, action and intent channels a policy needs from a human demonstration.

The constraint in robot learning is no longer video volume. It is the fraction of captured hours that pass a lab's quality bar, and the lag before they are labelled. Desmo is built to raise the first and remove the second.

Curation

Quality gate at capture time

Gaze stability and pupil response score every frame as it is recorded. Footage that would fail a hand-pose quality bar is flagged on the puck before it ships, not after annotation.

GazePupillometryOn-device compute
Labels

Measured, not estimated

Egocentric video at 30 fps underdetermines limb trajectory. Wrist trajectory comes from 200 Hz IMU instead, attended target from gaze, head pose from visual-inertial tracking. Every label is read off the sensors on one clock, with no video-estimation pipeline and no post-hoc alignment.

Wrist IMUHead IMUGaze
Annotation

Attention, annotated for free

The eye fixates a target several hundred milliseconds before the hand moves. Every session ships with gaze-grounded object labels: which object, from when, to when, with no annotation cost and no annotation lag.

GazeScene cam
Spatial

Pose without the alignment tax

Visual-inertial odometry needs IMU and video on a common time base. Both leave the recorder already stamped, so head pose and scene reconstruction need no temporal calibration step.

Head IMUScene cam
Delivery

Your schema, our export

Sessions are delivered as VRS, LeRobot, or RLDS, with per-file health checks. PII is blurred on-device before anything leaves the puck.

VRSLeRobotRLDSPII blur
Adjacent fields Same recording, same clock

Human factors

Workload, fatigue and attention lapses quantified during the task itself, in the cab, the cockpit or on the floor.

Teleoperation

Gaze and head pose as a control channel, logged alongside the scene the operator was working from.

Assistive AR

On-device context from scene video and gaze, with the compute to run inference at capture time rather than after the session.

04 / Research

Our claims come with citations.

Read more about our studies beyond the platform.

Abstract aperture figure, gaze stability and pupil novelty
arXiv:2603.04098 cs.CV · Mar 2026

Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning

Gaze confidence indexes visual stability; pupil response indexes information novelty. Gating on the first and ranking on the second retains 10% of frames at full-stream activity-recognition accuracy, with no model inference at capture time.

Read the paper →
Abstract stack figure, individual neural signatures across model depth
arXiv:2603.21847 cs.CL · Mar 2026

Riding Brainwaves in LLM Space: Understanding Activation Patterns Using Individual Neural Signatures

Frozen LLM representations contain person-specific neural directions. Per-participant linear probes on word-level EEG outperform a population probe ninefold on high-gamma power, and the directions do not transfer between individuals.

Read the paper →
05 / Design partners

Help shape Desmo.

We’re looking for a small group of research teams to design and validate Desmo with us in real-world studies. Tell us about your workflow, sensing needs, and what you want to learn—we’ll explore a partnership together.