Solutions / Video

Video of real people doing real tasks.

Human activity, gestures, interactions, and first-person task footage — captured to a named spec, with the framing, environments, and edge cases your model needs, not whatever happened to be lying around on the web.
A person wearing a small head-mounted camera on an elastic strap chops vegetables at a kitchen counter.
Egocentric, first-person task footage — our core specialism — captured while a contributor performs a real workflow.

Definition

Video data collection

Video data collection is the recording of moving-image footage from real contributors against a defined spec — activities, gestures, or first-person task views — with signed consent and controlled coverage of conditions.

What we collect

The formats teams ask us for.

Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.

Human activity & actionPeople performing specified actions and tasks for action recognition and activity understanding.
Gesture & interactionHand and body gestures, and human-to-human or human-to-device interaction sequences.
Egocentric / first-personHead-mounted task footage — our core specialism — with hands and manipulated objects in frame.
Multi-view & sceneFixed and moving third-person views of the same activity for cross-view training.
Domain workflowsReal workflows in kitchens, warehouses, retail, workshops, and clinical training environments.
Edge & failure casesOcclusion, motion, low light, and the transitions most datasets edit out.

What shapes the spec

The decisions we settle before collecting.

Viewpoint
Egocentric, third-person, or multi-view as specified
Environments
Real settings sampled to a coverage matrix, not a single staged room
Resolution & rate
Specified per project; egocentric capture uses our synchronized multi-sensor rig
Delivery
Video with per-clip metadata, timestamps, and optional action segments and labels
Consent basis
Written release per participant; bystanders de-identified or the clip discarded

Where it is used

What teams train with it.

  • Action and activity recognition
  • Gesture and interaction models
  • Embodied AI and robot manipulation (egocentric)
  • Sports, fitness, and physical-task analysis
  • Safety and procedure-compliance evaluation

How we run it

The standards behind every batch.

These apply to every modality — they are the reason the data is usable rather than merely large.

Collected to a written spec
Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
Contributors paid hourly
People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
Documented, informed consent
Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
Discard rather than downgrade
If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
Multi-layer QA with a visible reject log
Automated checks plus human review, and you see the reject reasons, not just the accepted items.
Buyer-owned commercial license
You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.

FAQ

Video collection, answered.

What is egocentric video, and do you collect it?
Egocentric video is first-person footage from a head-mounted camera, so the view matches what the wearer sees — hands entering frame, occlusion during grasp, motion during reach. It is our core specialism: we capture it on a synchronized multi-sensor rig with depth and hand pose. See our egocentric data catalog and capture methodology for the full detail.
Can you collect video of specific tasks or workflows?
Yes. We collect against a named spec — the exact actions, environments, viewpoints, and edge cases you need — rather than scraping generic clips. That is the difference between footage a model can learn from and footage teams discard.
How do you handle privacy in video?
Every participant signs a written release before capture, bystander faces and identifying details are removed, and any clip where a bystander cannot be de-identified is discarded. No covert capture, and no participants under 18.

Scope a video collection.

Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.