Solutions / Video
Video of real people doing real tasks.
Human activity, gestures, interactions, and first-person task footage — captured to a named spec, with the framing, environments, and edge cases your model needs, not whatever happened to be lying around on the web.

Definition
Video data collection
Video data collection is the recording of moving-image footage from real contributors against a defined spec — activities, gestures, or first-person task views — with signed consent and controlled coverage of conditions.
What we collect
The formats teams ask us for.
Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.
Human activity & actionPeople performing specified actions and tasks for action recognition and activity understanding.
Gesture & interactionHand and body gestures, and human-to-human or human-to-device interaction sequences.
Egocentric / first-personHead-mounted task footage — our core specialism — with hands and manipulated objects in frame.
Multi-view & sceneFixed and moving third-person views of the same activity for cross-view training.
Domain workflowsReal workflows in kitchens, warehouses, retail, workshops, and clinical training environments.
Edge & failure casesOcclusion, motion, low light, and the transitions most datasets edit out.
What shapes the spec
The decisions we settle before collecting.
- Viewpoint
- Egocentric, third-person, or multi-view as specified
- Environments
- Real settings sampled to a coverage matrix, not a single staged room
- Resolution & rate
- Specified per project; egocentric capture uses our synchronized multi-sensor rig
- Delivery
- Video with per-clip metadata, timestamps, and optional action segments and labels
- Consent basis
- Written release per participant; bystanders de-identified or the clip discarded
Where it is used
What teams train with it.
- Action and activity recognition
- Gesture and interaction models
- Embodied AI and robot manipulation (egocentric)
- Sports, fitness, and physical-task analysis
- Safety and procedure-compliance evaluation
How we run it
The standards behind every batch.
These apply to every modality — they are the reason the data is usable rather than merely large.
- Collected to a written spec
- Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
- Contributors paid hourly
- People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
- Documented, informed consent
- Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
- Discard rather than downgrade
- If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
- Multi-layer QA with a visible reject log
- Automated checks plus human review, and you see the reject reasons, not just the accepted items.
- Buyer-owned commercial license
- You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.
FAQ
Video collection, answered.
- What is egocentric video, and do you collect it?
- Egocentric video is first-person footage from a head-mounted camera, so the view matches what the wearer sees — hands entering frame, occlusion during grasp, motion during reach. It is our core specialism: we capture it on a synchronized multi-sensor rig with depth and hand pose. See our egocentric data catalog and capture methodology for the full detail.
- Can you collect video of specific tasks or workflows?
- Yes. We collect against a named spec — the exact actions, environments, viewpoints, and edge cases you need — rather than scraping generic clips. That is the difference between footage a model can learn from and footage teams discard.
- How do you handle privacy in video?
- Every participant signs a written release before capture, bystander faces and identifying details are removed, and any clip where a bystander cannot be de-identified is discarded. No covert capture, and no participants under 18.
Scope a video collection.
Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.