Natural English Dialogue Sets
Natural multi-speaker conversation with the turn structure, timing and speaker context needed for ASR, diarisation and conversational systems.

Byser designs and delivers purpose-built datasets for AI training and evaluation, with the structure, annotations and metadata your model task requires.
Model task
The same raw media can become very different datasets depending on what a model needs to learn. Byser defines the useful unit of data, the relationships between fields and the technical context around each example, so the collection reflects the job it is meant to do.
Transcripts, OCR, temporal segments, labels and other supervisory signals can be attached where they add useful information to the task.
Source and label provenance can remain visible in the data, helping technical teams distinguish different kinds of supervision instead of treating every field as equivalent.
Media, records and metadata can be connected through stable identifiers and machine-readable manifests, reducing the work required to understand how the package fits together.
Dataset collections
Speech, screen, video, audio, image and multimodal formats, each designed around a distinct modelling problem.
Natural multi-speaker conversation with the turn structure, timing and speaker context needed for ASR, diarisation and conversational systems.
Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.
Speech organised across defined accent and regional groups for adaptation, coverage analysis and comparative model evaluation.
Speech data for languages and varieties where existing machine-learning resources are limited, fragmented or difficult to standardise.
Spoken instructions and short follow-up exchanges that preserve intent, correction, confirmation and clarification behaviour.
Bilingual conversation in which language changes can be represented within turns rather than reduced to a single session label.
Start with the model problem. We can define the modality, unit of observation, annotations, metadata and delivery structure around the way the data will actually be used.
Technical delivery
Depending on the collection, delivery can include task-specific records, machine-readable manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions. The point is not to add more files. It is to make the relationship between the source data and the model task clear.
Share the task, the data conditions that matter and the structure you need to work with. We will use that to define the collection.