First-Person Activity Clips
Egocentric activity video for models that need to understand actions, objects and task progression from the actor's point of view.
View datasetManipulation is easier to model when an action is connected to the object it affects and the state change that follows. Hand and Object Interaction Clips focus on that relationship at close range rather than treating the whole video as one global action label.
The collection can localise interaction intervals and keep object identity consistent across a sequence, making it possible to represent contact, movement and before-and-after state as related parts of the same example.
Manipulation learning, fine-grained action recognition, grasp modelling, state-change prediction and robotics research.
The action vocabulary and state representation are defined around the objects and manipulation task selected for the collection.
The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.