Cross-Language Dialogue Sets
Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.
View datasetCode-switching is not captured well by labelling an entire recording with two languages. The useful information often lies in where the language changes, how long each span lasts and what happens in the conversation around the switch.
Bilingual Code-Switching Dialogue Sets can represent language at the span or turn level while keeping speaker and timing information connected to the same exchange. That gives speech and language models a clearer signal for the transition itself rather than only the fact that two languages are present somewhere in the session.
Code-switched ASR, language identification, switch-point detection and multilingual dialogue systems.
Language pairs and span granularity are defined according to the behaviour the model needs to learn or measure.
The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.