Skip to main content
AUDIOLANGUAGE COVERAGE

Rare and Low-Resource Language Sets

For low-resource languages, the difficulty is often not only finding speech. Scripts, spelling conventions, language variants and transcript practices may differ enough that a generic speech schema creates more ambiguity than it removes.

Rare and Low-Resource Language Sets are designed to preserve those language-specific requirements inside a consistent delivery structure. The collection can connect recordings to native-script transcripts and the metadata needed to understand which variety, orthographic convention or additional representation is being used.

Technical structure

  • Native-script transcripts with language and variant identifiers
  • Script and orthographic metadata where useful to the task
  • Optional romanisation, translation or pronunciation-related fields
  • Shared record structure across the collection without flattening language-specific conventions

Typical applications

Low-resource ASR, language identification, multilingual expansion and language-specific speech modelling.

Delivery note

The transcript policy and language metadata are defined around the linguistic reality of the target language or variety.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections