Skip to main content
VIDEOEGOCENTRIC

First-Person Activity Clips

First-person video changes the learning problem because the camera sees the activity from the actor's point of view. Hands, nearby objects and task progression appear in a perspective closer to the input available to an embodied system.

First-Person Activity Clips can preserve that egocentric sequence while adding temporal structure around actions or task phases. When speech or narration matters, the audio and transcript can also be connected to the relevant part of the activity instead of being treated as unrelated metadata.

Technical structure

  • Egocentric video linked to activity and session identifiers
  • Temporal segments for actions or task phases where useful
  • Object, environment and capture metadata relevant to the scenario
  • Optional narration, transcript or audio relationships when speech contributes to the task

Typical applications

Egocentric activity recognition, embodied AI, robotics perception, temporal action understanding and procedural modelling.

Delivery note

Activities, camera configuration and temporal granularity are selected around the intended perception or action task.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections