Skip to main content
AUDIOCOMMAND SPEECH

Instruction and Command Voice Sessions

Voice systems rarely operate as a list of isolated trigger phrases. A user gives an instruction, corrects a detail, confirms an interpretation or clarifies what they meant. Instruction and Command Voice Sessions preserve that short interaction around the command itself.

The resulting records can connect spoken language to intent, slots and turn relationships, giving models a clearer representation of how commands evolve across a brief exchange. This is especially useful when the system must recover from ambiguity rather than simply recognise a keyword.

Technical structure

  • Command and follow-up turns linked within a session
  • Intent and slot structures defined around the use case
  • Correction, clarification and confirmation relationships where relevant
  • Device, environment or speaker metadata where those conditions affect the task

Typical applications

Voice agents, intent classification, slot filling, spoken instruction following and command recovery.

Delivery note

The intent taxonomy and session design are shaped around the command behaviour the model needs to handle.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections