Cross-Language Dialogue Sets
Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.
View datasetFor low-resource languages, the difficulty is often not only finding speech. Scripts, spelling conventions, language variants and transcript practices may differ enough that a generic speech schema creates more ambiguity than it removes.
Rare and Low-Resource Language Sets are designed to preserve those language-specific requirements inside a consistent delivery structure. The collection can connect recordings to native-script transcripts and the metadata needed to understand which variety, orthographic convention or additional representation is being used.
Low-resource ASR, language identification, multilingual expansion and language-specific speech modelling.
The transcript policy and language metadata are defined around the linguistic reality of the target language or variety.
The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.