Skip to main content
VIDEOMANIPULATION

Hand and Object Interaction Clips

Manipulation is easier to model when an action is connected to the object it affects and the state change that follows. Hand and Object Interaction Clips focus on that relationship at close range rather than treating the whole video as one global action label.

The collection can localise interaction intervals and keep object identity consistent across a sequence, making it possible to represent contact, movement and before-and-after state as related parts of the same example.

Technical structure

  • Object identifiers linked across the interaction sequence
  • Action phases and temporal boundaries where useful
  • Hand-object contact or manipulation information where relevant
  • Object state before and after the action when the task depends on it

Typical applications

Manipulation learning, fine-grained action recognition, grasp modelling, state-change prediction and robotics research.

Delivery note

The action vocabulary and state representation are defined around the objects and manipulation task selected for the collection.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections