Skip to main content
AUDIOCODE-SWITCHING

Bilingual Code-Switching Dialogue Sets

Code-switching is not captured well by labelling an entire recording with two languages. The useful information often lies in where the language changes, how long each span lasts and what happens in the conversation around the switch.

Bilingual Code-Switching Dialogue Sets can represent language at the span or turn level while keeping speaker and timing information connected to the same exchange. That gives speech and language models a clearer signal for the transition itself rather than only the fact that two languages are present somewhere in the session.

Technical structure

  • Bilingual transcripts with language-span or turn-level language labels
  • Switch boundaries linked to utterance timing
  • Speaker-turn structure retained across the conversation
  • Optional translation or language-pair metadata where relevant

Typical applications

Code-switched ASR, language identification, switch-point detection and multilingual dialogue systems.

Delivery note

Language pairs and span granularity are defined according to the behaviour the model needs to learn or measure.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections