Egocentric capture · embodied AI

First-person video that survives contact with a policy.

Egocentric capture for embodied AI. Every clip ships with depth, hand pose, and body pose — synchronized to under two milliseconds, with a documented consent chain and a commercial license attached.

Sync ceiling
2.0 ms, hardware-triggered
License
Commercial, buyer-owned
Enrichment
Depth · hands · body · masks
Delivery
RLDS · LeRobot · WebDataset
bowl 0.97cutting_board 0.94mug 0.99bottle 0.91WRIST_RWRIST_LBODY 24 JOINTS · 3DHANDS 21 × 2 · 3D METRICGAZE · mugREC 00:04:18:163840×2160 · 60 FPSSYNC Δ 1.5 msEP_0117_KITCHEN_A / rig-04
00:04:18:16054/300 f

Synthetic frame, real schema. Toggle any layer — every one of these ships on every clip.

1,247hours

Validated hours delivered

how we count

Hours that passed sync + QA and shipped to a buyer. Recaptured and rejected hours excluded.

1.12ms

Median cross-stream sync error

how we count

Median absolute timestamp offset between head camera and every other stream, over 12,480 measured frames.

100%

Documented consent

how we count

Every participant signs a written commercial release before capture. Bystanders are de-identified or the clip is discarded.

9days

Median time to first delivery

how we count

Signed spec to first validated batch in your hands. Measured across Spec Pilot engagements.

The sample

See the data before you talk to anyone.

A 30-second slice with nothing stripped out: same schema, same enrichment layers, same license terms as a production batch. No form, no call, no NDA.

Rendered from the schema below. Toggle layers, scrub the playhead, then read the files it came from.

bowl 0.97cutting_board 0.94mug 0.99bottle 0.91WRIST_RWRIST_LBODY 24 JOINTS · 3DHANDS 21 × 2 · 3D METRICGAZE · mugREC 00:04:18:163840×2160 · 60 FPSSYNC Δ 1.5 msEP_0117_KITCHEN_A / rig-04
00:04:18:16054/300 f

Download the full 2.4 GB pack

40 episodes across four environments, all streams at native rate. Takes an email — asked for after the free files above, not before.

What ships with every clip

Ten streams, one clock, one episode file.

Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.

StreamSpecDetail
Head camera3840 × 2160 · 60 fpsGlobal-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes
Wrist cameras × 21920 × 1080 · 60 fpsLeft + right, 100° HFOV, rigid mount, extrinsics re-solved per session
Depth848 × 480 · 30 fpsActive stereo, 0.3–4 m range, metric millimetres, per-frame confidence map
IMU200 Hz · 6-DoFAccel + gyro, bias-calibrated, hardware-timestamped on the same clock domain
Hand pose21 keypoints × 2 hands3D metric, per-joint visibility flag, contact events on grasp and release
Body pose24 joints3D, root-relative and world-frame, torso and forearm chains resolved
SegmentationInstance masksManipulated objects + target surfaces, tracked IDs across the episode
Action segmentsVerb + noun taxonomy97 verbs, 512 nouns, start/end to the frame, human-reviewed
FormatsRLDS · LeRobot · WebDatasetAlso HDF5, zarr, and .rrd for Rerun. Converters shipped as source.
LicenseCommercial · buyer-ownedPerpetual, irrevocable, model-weights-clean. Exclusivity available.

Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.

How we capture

One clock domain, or it isn't a dataset.

Ten streams recorded on separate sensors are ten separate videos until something forces them onto the same timebase. We hardware-trigger every sensor from a single PTP grandmaster, and we publish the resulting error distribution.

A head-mounted capture rig disassembled and laid out: head camera, two wrist cameras, depth sensor, IMU pucks, cables and battery pack.
FIG 1 — Rig v3, 640 g total worn weight. Every part is off-the-shelf and named below, so you can reproduce our capture conditions.
A person wearing the head-mounted rig while chopping vegetables in a real, cluttered home kitchen.
Bill of materials
Head cameraSony IMX585 global shutter · 120° M12Carbon headband, 42 g
Wrist camerasIMX290 ×2 · 100° M12Velcro cuff, rigid plate
DepthIntel RealSense D435ifForehead, 18 mm below head cam
IMUBosch BMI088 ×3 · 200 HzHead, both wrists
ClockPTP grandmaster + hardware triggerBelt pack, 190 g
Storage2 TB NVMe · 6 h continuousBelt pack

Rig architecture

Four device groups, one trigger, one episode file. The clock node is the only part of this diagram that is hard.

Rig architecturerig-04 · v3 payload
HEAD CAM3840×2160 · 60 fpsglobal shutter · 120° HFOVWRIST CAM ×21920×1080 · 60 fpsrigid mount · 100° HFOVDEPTH848×480 · 30 fpsactive stereo · 0.3–4 mIMU200 Hz · 6-DoFbias-calibrated accel + gyroPTP CLOCKhardware triggerΔ < 2 ms across allone clock domain · no post-hoc alignUNIFIED EPISODEtimeStampNs (int64)cameraIntrinsics / extrinsicshandJoints[2][21][3]bodyJoints[24][3]instanceMasks[objects]actionSegments[verb, noun]RLDS · LeRobot · WebDataset · HDF5 · zarr · .rrd

Measured sync error

Absolute timestamp offset between the head camera and every other stream, sampled across randomly selected frames from shipped episodes.

Cross-stream sync errorn = 12,480 frames
0%25%50%75%100%SPEC CEILING 2.0 msMEDIAN 1.12 msp99 1.940.01.02.03.04.0ABSOLUTE OFFSET FROM HEAD CAMERA (ms)
12,480 measured frames. p99 = 1.94 ms. Frames above the ceiling are rejected and the segment is recaptured.

Method, tooling and the raw CSV behind this plot are in the sync measurement guide. Nobody else in this category publishes this number, which is exactly why we do.

A checkerboard calibration target on a tripod facing a camera rig in a plain room.
FIG 2 — Intrinsics and extrinsics are re-solved at the start and end of every session. Drift between the two solves is recorded in the episode metadata, so it is detectable rather than assumed.

Coverage — and what we don't have

Everyone publishes a bigger number. We publish the histogram, including the gaps.

If your task lives in a thin column, that is a spec conversation, not a no. Thin cells are where we are actively accepting specs, and where a pilot buys you exclusivity on the capture.

Validated hours by scenarioas of 2026-07
01k2k3kKitchen & food prep2,840Household chores2,310Tool use & assembly1,680Warehouse & logistics1,420Retail & shelf ops890Automotive service640Construction & trades410Clinical & care180Agriculture120Hospitality95VALIDATED HOURS
DEEP — SHIPPING NOWTHIN — ACCEPTING SPEC
Coverage matrix · environment × conditionindex 0–100
DaylightMixedLow lightClutteredMulti-personFailure cases94886179442890825571512272763866581968713362641477644183293158421247369KitchenLiving spaceWarehouseRetail floorWorkshopOutdoor site
Index = validated hours in that cell, normalised against our deepest cell. Failure cases are deliberate mistakes, drops, and recoveries — the hardest thing to source and the column we are shortest on.

Evidence it trains better

The only number that matters is the one downstream of the data.

Same architecture, same eval suite, same seeds — only the pretraining corpus changes. Hours are on the label so you can see we are not winning on volume.

Downstream success · 40 unseen manipulation tasksidentical architecture · identical eval suite
0%20%40%60%80%Ego4D3,670 h31%EgoDex829 h38%EgoScale20,854 h54%Firsthand1,200 h67%SUCCESS RATE, 40 UNSEEN TASKS
ILLUSTRATIVE — REPLACE WITH MEASURED RESULTS BEFORE PUBLISHINGbenchmark repo →

Training code, eval harness and per-task breakdown live in the benchmark repo. If a number here does not reproduce on your cluster, we want the issue filed.

Architecture
Identical diffusion policy, 210 M params, no per-dataset tuning
Eval suite
40 unseen manipulation tasks, 20 trials each, fixed seeds
Success criterion
Task-completion predicate, scored by held-out rubric, not by us
Confound
Firsthand is 1,200 h vs EgoScale 20,854 h — the delta is fitness, not volume

Provenance and consent

The license is a product surface, not a footnote.

Your legal team will read this section before your research team reads the specs. So here is the actual release language, the actual pay, and the actual de-identification pipeline.

Participant release — verbatim

PARTICIPANT RELEASE — COMMERCIAL TRAINING USE
Version 3.2 · effective 2026-01-14

1. GRANT. I grant Firsthand Data, Inc. ("Firsthand") a perpetual,
   irrevocable, worldwide, sublicensable, royalty-free right to
   record, store, process, reproduce, and distribute the audio,
   video, depth, and motion recordings captured of me during the
   session(s) described in Schedule A (the "Recordings"), and to
   license the Recordings to third parties for the purpose of
   training, evaluating, and deploying machine learning models,
   including commercial models.

2. SCOPE. The grant covers derivative representations of the
   Recordings, including but not limited to depth maps, hand and
   body pose estimates, segmentation masks, textual action labels,
   and model weights derived from any of the foregoing.

3. COMPENSATION. I will be paid the hourly rate stated in
   Schedule A for all time spent in session, including setup,
   calibration, and breaks. Payment is not contingent on the
   Recordings being accepted, delivered, or licensed.

4. WITHDRAWAL. I may withdraw consent for future sessions at any
   time. I may request deletion of any Recording not yet delivered
   to a licensee by written notice to privacy@ifirsthand.com.
   Recordings already delivered cannot be recalled from a licensee,
   and I acknowledge this limitation.

5. BYSTANDERS. I confirm that I have not knowingly recorded any
   person who has not signed this release. Where an unidentified
   person appears, Firsthand will de-identify them or discard the
   affected footage.

6. NO IDENTITY LICENSE. This release does not grant the right to
   use my name, likeness, or voice for advertising or endorsement.

7. GOVERNING LAW. Delaware, USA, without regard to conflict of
   laws principles.

Participant signature: ______________________  Date: __________
Operator signature:    ______________________  Date: __________

What contributors are paid

Base session rate
$38 / hour
Includes
Setup, calibration, breaks, travel over 20 min
Thin-coverage premium
+$12 / hour (clinical, construction, agriculture)
Payment terms
Net 7, no per-task bounty, no rejection clawback

Published because a per-task bounty produces rushed, unusable footage. Collector standards.

De-identification pipeline

Faces
Bystander faces detected and blurred; participant faces mostly out of frame by rig geometry
Plates
Licence plates and street numbers blurred on all outdoor and automotive footage
Audio
Speech scrubbed of names, addresses and card numbers; ambient audio retained
Screens
Visible phone and monitor content masked unless it is the manipulated object
Bystanders
If a person cannot be de-identified, the clip is discarded rather than shipped

Jurisdictions of capture

  • United States
  • Canada
  • United Kingdom
  • Germany
  • Netherlands
  • Poland
  • Portugal
  • Japan
  • South Korea
  • Singapore
  • Mexico
  • Brazil
  • GDPR ART. 6(1)(a)
  • CCPA / CPRA
  • UK GDPR
  • APPI (JP)
  • PIPEDA (CA)
  • SOC 2 TYPE II — IN AUDIT

Case study

A bimanual loading policy that kept failing on half-full racks.

Anonymised at the buyer's request: a US humanoid lab, Series B, training a bimanual dishwasher-loading policy. Their footage was clean and their policy still failed the moment a rack was partially occupied.

620hours
Validated hours delivered in 11 weeks
11skills
Named skills in the spec, all cells filled
+21pts
Success rate on their internal dishwasher eval
7.4%
Batch reject rate — recaptured at our cost

The lab had 4,000 hours of public egocentric footage and a policy that scored well in simulation. In the real world it stalled on any rack that was already half loaded, because almost nothing in the public corpora shows a person recovering from a bad grasp inside a cluttered fixture.

We wrote a spec with a hard quota: 15% of every batch had to be a failure and recovery — dropped plate, wrong slot, tine collision — with the recovery captured through to completion. That quota is the whole case study.

Their words, lightly trimmed: “The failure cases were the only part we couldn’t buy anywhere else, and they moved the eval more than the other 500 hours combined.”

A dim multi-monitor QA workstation showing synchronized video tracks on a timeline next to a spreadsheet of measurements.
FIG 3 — QA station. Every batch is reviewed frame-accurate against the accept criteria before it ships. Rejects and their reasons go to the buyer with the batch.

Engagement timeline

  1. WEEK 0Spec written together: 11 skills, 6 kitchens, explicit failure-case quota of 15%
  2. WEEK 1Rig calibration on their object set — 34 SKUs of real crockery shipped to us
  3. WEEK 2First 50 h batch delivered in RLDS. They found a wrist extrinsics error. We recaptured 12 h.
  4. WEEK 4Cadence at 70 h/week. Coverage report per batch with the reject log attached.
  5. WEEK 11620 h delivered. Failure cases turned out to be the highest-value slice.

Anonymised with permission. Reference call available under NDA once a spec is scoped.

Pipeline

Five steps. Spec to delivery.

Step 03 is the one most vendors leave out, and it is the reason you are not paying for footage you will throw away.

  1. 01SPECNamed skill list, coverage targets, accept/reject criteria
  2. 02CAPTURECalibrated rigs, vetted and trained operators, real environments
  3. 03VALIDATESync check, QA pass, reject and recapture — you see the reject log
  4. 04ENRICHDepth, body pose, hand pose, action segments, instance masks
  5. 05DELIVERRLDS, LeRobot, WebDataset, HDF5 — plus the data card

FAQ

Questions we get on the first call.

Answered here so the first call can be about your skill spec instead.

01What is egocentric video data?

Egocentric video is footage recorded from a head-mounted camera worn by the person performing a task, so the camera shares the actor’s viewpoint and moves with their head. For embodied AI it matters because the observation distribution matches what a robot with a head or chest camera actually sees: hands entering frame from below, occlusion during grasp, motion blur during reach, and a first-person view of the manipulated object. Third-person footage does not contain that geometry, so policies trained on it transfer poorly.

02How much data do I need to train an embodied AI model?

NVIDIA’s EgoScale work reports a log-linear relationship between egocentric hours and downstream policy success out to 20,854 hours with no observed saturation, so more hours keep helping. In practice the binding constraint is fitness, not volume: teams typically discard around 90% of bulk footage as unusable for manipulation. A useful starting point is 50–100 validated hours against one named skill, which is enough to measure whether your architecture responds to the data before you commit to a thousand hours.

03How is this different from Ego4D or EgoDex?

Ego4D’s license does permit commercial use, so the honest gap is fitness rather than legality: it is roughly 3,670 hours of unstructured daily activity, much of it with heavy motion blur, no manipulation-grade 3D hand pose, and no frame-accurate action labels. EgoDex has excellent hand pose but is licensed CC BY-NC-ND, which genuinely rules out commercial training use. Firsthand captures against a named skill spec, hardware-synchronizes every stream to under two milliseconds, ships depth, 3D hand pose, body pose, segmentation and action segments on every clip, and attaches a commercial buyer-owned license with a documented consent chain.

04What hardware do you capture on?

A head-mounted global-shutter 4K/60 camera with a 120° lens, two wrist-mounted 1080p/60 cameras, an active-stereo depth sensor at 848×480/30, and 200 Hz 6-DoF IMUs. All streams are hardware-triggered onto a single PTP clock domain, and intrinsics and extrinsics are re-solved with a checkerboard target at the start and end of every session so drift is detectable rather than assumed.

05What does it cost?

We quote per engagement rather than publishing a rate card, because the price is driven by things we cannot guess from a web page: which environments you need, how thin the coverage is in them, how dense the annotation has to be, and whether you want exclusivity. The sample pack is free and ungated, so you can evaluate the data before any commercial conversation. Tell us what you are training and your target hours in the intake form and you will get a scoped estimate — hours, price, delivery date, and the coverage cells we would need to fill — within one business day. Billing is always per validated hour: footage that passed sync and QA against your spec. Rejected episodes, recaptures, calibration time and travel are never billed.

06How do you handle consent and licensing?

Every participant signs a written release granting commercial training rights before capture begins, and is paid a published hourly rate rather than a per-task bounty. Faces and licence plates of bystanders are blurred, audio is scrubbed of identifying speech, and any clip where a bystander cannot be de-identified is discarded rather than shipped. You receive a perpetual, irrevocable, buyer-owned commercial license, the consent artefacts for the batch, and a data card listing jurisdictions of capture.

07What formats do you deliver in?

RLDS, LeRobot, WebDataset, HDF5, zarr, and Rerun .rrd. Every episode carries int64 nanosecond timestamps per stream, camera intrinsics and extrinsics, handJoints[2][21][3], bodyJoints[24][3], instance masks, and action segments. Converters ship as readable source, not a binary, so you can retarget the schema to your own loader.

08How fast can you start?

Spec conversation the same week, first validated batch in a median of nine days from a signed spec. Thin coverage areas — construction, clinical, agriculture, hospitality — take longer to staff, typically three to four weeks to first batch, because operator vetting in those environments is the bottleneck.

Request a quote

Tell us what you're training.

We quote per engagement rather than off a rate card, because price depends on which environments you need, how thin our coverage is in them, and how dense the annotation has to be. Fill this in and you get real numbers back.

Reply time
One business day, with a scoped estimate
You receive
Hours, price, delivery date, coverage cells to fill
First call
30 minutes, technical, no deck
Then
A written spec you own, whether or not you buy
Billing unit
Validated hours only — rejects are never billed

Or skip the form entirely — hello@ifirsthand.com. A research engineer reads it, not a sales inbox.