Cross-Language Dialogue Sets
Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.
View datasetNatural conversation contains signals that disappear when speech is reduced to isolated sentences. Speakers interrupt, hesitate, overlap, correct themselves and respond to what happened several turns earlier. Natural English Dialogue Sets preserve those relationships so the data reflects how conversation actually unfolds.
A session can connect the audio to speaker-labelled transcripts, turn boundaries and utterance-level timing, with additional dialogue information added where it serves the task. For ASR and conversational systems, this makes the supervision more useful than a transcript alone: the model can work with who spoke, when the turn changed and how the exchange developed over time.
Conversational ASR, diarisation, turn-taking, voice agents, dialogue modelling and conversation analysis.
The final schema is shaped around the transcript depth, speaker structure and dialogue signals required by the project.
The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.