Solutions / Multimodal & sensor

Multiple streams, captured on one clock.

Synchronized combinations of video, audio, depth, motion, and pose — for robotics, embodied AI, and multimodal models that need the streams to line up, not just coexist in the same folder.
A multi-sensor head-mounted capture rig on a stand, with cameras, a depth sensor, and cabling, against a dark studio background.
Our synchronized egocentric rig — head and wrist cameras, depth, IMU, and pose triggered onto a single clock, verified under two milliseconds.

Definition

Multimodal data collection

Multimodal data collection is the simultaneous capture of two or more aligned data streams — for example video, depth, audio, and motion — from the same event, time-synchronized so a model can learn the relationships between them.

What we collect

The formats teams ask us for.

Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.

Synchronized video + depthRGB and metric depth aligned per frame for 3D perception and manipulation.
Motion & IMUAccelerometer and gyroscope streams, hardware-timestamped on the same clock as video.
Hand & body pose3D hand and body keypoints tracked through the sequence, with visibility and contact events.
Audio + visualAligned speech or event audio with video for audio-visual models.
Sensor & robotics dataMulti-sensor rigs for robotics and IoT, aligned to a single timeline.
Egocentric multi-streamOur core rig: head and wrist cameras, depth, IMU, and pose synchronized under two milliseconds.

What shapes the spec

The decisions we settle before collecting.

Synchronization
Streams hardware-triggered onto a shared clock; egocentric rig verified under 2 ms
Stream set
Chosen per project — video, depth, audio, IMU, pose, segmentation
Calibration
Intrinsics and extrinsics re-solved per session so drift is detectable, not assumed
Delivery
RLDS, LeRobot, WebDataset, HDF5, zarr, or Rerun .rrd, with per-stream timestamps
Consent basis
Written release per participant; discard-don’t-downgrade on de-identification

Where it is used

What teams train with it.

  • Robot manipulation and imitation learning
  • Embodied AI and world models
  • Audio-visual understanding
  • Activity recognition from fused sensors
  • IoT and wearable sensor models

How we run it

The standards behind every batch.

These apply to every modality — they are the reason the data is usable rather than merely large.

Collected to a written spec
Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
Contributors paid hourly
People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
Documented, informed consent
Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
Discard rather than downgrade
If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
Multi-layer QA with a visible reject log
Automated checks plus human review, and you see the reject reasons, not just the accepted items.
Buyer-owned commercial license
You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.

FAQ

Multimodal & sensor collection, answered.

What does multimodal data collection actually involve?
Capturing two or more streams from the same event — say video, depth, audio, and motion — and aligning them on a single clock so a model can learn how they relate. The hard part is synchronization: streams that are merely in the same folder but drift against each other are close to useless for manipulation. We hardware-trigger onto a shared clock and verify the offset.
How tightly are the streams synchronized?
For our egocentric rig, every stream is hardware-triggered onto one PTP clock domain and the median cross-stream offset is verified under two milliseconds. For other sensor combinations we agree and measure the synchronization target in the spec, and report the achieved offset with the batch.
What formats do you deliver multimodal data in?
RLDS, LeRobot, WebDataset, HDF5, zarr, and Rerun .rrd, with int64 nanosecond timestamps per stream and calibration included. Converters ship as readable source so you can retarget to your own loader.

Scope a multimodal & sensor collection.

Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.