Natural English Dialogue Sets
Natural multi-speaker conversation with the turn structure, timing and speaker context needed for ASR, diarisation and conversational systems.
Catalogue
Purpose-built formats for AI training and evaluation across speech, screen, video, audio, image and multimodal data.
Speech and dialogue
Natural multi-speaker conversation with the turn structure, timing and speaker context needed for ASR, diarisation and conversational systems.
Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.
Speech organised across defined accent and regional groups for adaptation, coverage analysis and comparative model evaluation.
Speech data for languages and varieties where existing machine-learning resources are limited, fragmented or difficult to standardise.
Spoken instructions and short follow-up exchanges that preserve intent, correction, confirmation and clarification behaviour.
Bilingual conversation in which language changes can be represented within turns rather than reduced to a single session label.
Speech captured under defined delivery styles for work on prosody, pacing, emphasis and style-conditioned modelling.
Scenario-based support conversations that preserve roles, changing intent, escalation and the outcome of the exchange.
Screen and software
Desktop task recordings that connect interface state, user action and task outcome across a continuous workflow.
Browser-based task sequences for models that need to reason across navigation, page state and multi-step interaction.
Screen workflows centred on failed actions, visible error states, corrections and the path back towards task completion.
Mobile task sequences with device context and touch interaction where available, built for models that operate inside app interfaces.
Narrated software workflows that connect spoken explanation with screen state, visible text and user action.
Video
Egocentric activity video for models that need to understand actions, objects and task progression from the actor's point of view.
Close-range manipulation video focused on hand-object contact, fine-grained actions and changes in object state.
Multi-step demonstrations structured around procedure order, temporal boundaries, progress and completion.
Activity video captured across defined settings or conditions for robustness, domain-shift and context-sensitive model work.
Audio and multimodal
Ambient audio organised around acoustic scenes, recording conditions and sound events where the task requires them.
Linked audio, video, screen, image or text records with explicit relationships between modalities.
A collection designed from your own objective, schema, inputs, targets and delivery requirements rather than a predefined format.
Image
Product imagery across views, packaging surfaces and visual conditions for recognition, retrieval and packaging understanding.
Object images across viewpoints and states, with persistent instance relationships where the task depends on them.
Text-bearing scenes with explicit links between image regions, transcription, language and script for OCR and scene-text models.
If the task does not fit an existing format, we can define a collection around your own data contract or model requirement.