Shipping · 890 validated hours

Retail shelf stocking egocentric dataset

Facing, restocking, price changes and planogram resets on live retail floors with bystanders present and de-identified.
Episodes
6,210
Participants
96
Capture sites
31
Sync ceiling
2.0 ms, hardware-triggered

What a frame contains

Every layer, on every clip in this set.

Rendered from the delivered schema rather than a marketing composite. Toggle layers to see what arrives with the footage.

bowl 0.97cutting_board 0.94mug 0.99bottle 0.91WRIST_RWRIST_LBODY 24 JOINTS · 3DHANDS 21 × 2 · 3D METRICGAZE · mugREC 00:04:18:163840×2160 · 60 FPSSYNC Δ 1.5 msEP_0117_KITCHEN_A / rig-04
00:04:18:16054/300 f

Synthetic frame, real schema. The sample pack contains the same fields for actual episodes.

A grocery store shelf aisle with stocked shelves and a part-emptied case on the floor
Representative capture environment for this set. Rooms are recorded as found — clutter is signal, not something we tidy away before rolling.

Named skills

What the operators were told to do.

Coverage is specified per skill, not per hour. These are the skill lines currently filled in this environment.

  • Face and front stock to a planogram
  • Case-cut and shelf-load
  • Price and label swap
  • Chilled-case rotation by date code
  • Hanger and fixture reset

Environments captured

  • Grocery aisle
  • Chilled case
  • Apparel fixture
  • Back-of-house staging

Condition cells filled

Daylight
68%
Mixed light
71%
Multi-person
64% — customers in frame throughout
Failure cases
14% — the thinnest failure coverage in the catalog

Known failure modes

What goes wrong in this environment.

Published because you will find these in the data within an afternoon, and it is cheaper for both of us if you find them in this list first.

  • Every bystander is blurred at ingest; clips where a bystander cannot be de-identified are discarded, which is why episode yield here is lower than kitchen.
  • Chilled-case capture is time-boxed to protect stock, so episodes are shorter — median 41 s versus 78 s elsewhere.

What ships with every clip

Ten streams, one clock, one episode file.

Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.

StreamSpecDetail
Head camera3840 × 2160 · 60 fpsGlobal-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes
Wrist cameras × 21920 × 1080 · 60 fpsLeft + right, 100° HFOV, rigid mount, extrinsics re-solved per session
Depth848 × 480 · 30 fpsActive stereo, 0.3–4 m range, metric millimetres, per-frame confidence map
IMU200 Hz · 6-DoFAccel + gyro, bias-calibrated, hardware-timestamped on the same clock domain
Hand pose21 keypoints × 2 hands3D metric, per-joint visibility flag, contact events on grasp and release
Body pose24 joints3D, root-relative and world-frame, torso and forearm chains resolved
SegmentationInstance masksManipulated objects + target surfaces, tracked IDs across the episode
Action segmentsVerb + noun taxonomy97 verbs, 512 nouns, start/end to the frame, human-reviewed
FormatsRLDS · LeRobot · WebDatasetAlso HDF5, zarr, and .rrd for Rerun. Converters shipped as source.
LicenseCommercial · buyer-ownedPerpetual, irrevocable, model-weights-clean. Exclusivity available.

Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.

Data card

The card that ships with the batch.

Delivered as machine-readable JSON alongside the episodes, so provenance travels with the data instead of living in an email thread.

Dataset
Retail shelf stocking egocentric dataset
Validated hours
890 h — passed sync and QA, shipped to at least one buyer
Episodes / participants / sites
6,210 / 96 / 31
Sync ceiling
Δ < 2.0 ms across all streams · median 1.12 ms · p99 1.94 ms
Consent
Written commercial release per participant, signed before capture begins
De-identification
Faces and licence plates blurred, audio scrubbed, un-blurrable clips discarded
Collection period
Rolling. Batch capture dates recorded per episode.
License
Perpetual, irrevocable, buyer-owned commercial. Exclusivity available.
Known limitations
Every bystander is blurred at ingest; clips where a bystander cannot be de-identified are discarded, which is why episode yield here is lower than kitchen.
Formats
RLDS, LeRobot, WebDataset, HDF5, zarr, Rerun .rrd

Flagged frames are shipped rather than silently dropped. You decide whether to mask, down-weight, or exclude them.

Nearest public dataset

How this compares to MMAct / retail CCTV corpora.

Where a public set is genuinely better at something, we say so. Where the blocker is the license rather than the quality, we say that too.

MMAct / retail CCTV corpora

Hours
Varies
License
Research-only, mostly third-person

Existing retail footage is overwhelmingly fixed-camera surveillance, which gives you none of the first-person reach geometry. This is head- and wrist-mounted capture from the person doing the stocking.

Firsthand — Retail & shelf ops

Hours
890 h
License
Commercial, buyer-owned, perpetual

Captured against the named skill list above, hardware-synchronized under 2.0 ms, and shipped with depth, 3D hand pose, body pose, instance masks and action segments on every clip.

See the downstream result →

Pilot retail & shelf ops against your spec.

Write the skill list with us, get fifty to a hundred validated hours in two weeks, and read the reject log before you commit to volume.