Skip to main content
IMAGESCENE TEXT

Real-World Text and Signage Photos

Scene-text models need more than an image containing words. They need to know where the text appears, what it says and which language or script it uses. Real-World Text and Signage Photos make those relationships explicit at the region level.

Text geometry can be connected to transcription and scene context, with reading order or translation added where the task benefits from them. This creates a direct image-to-region-to-text structure for models working outside the clean conditions of scanned documents.

Technical structure

  • Text-region geometry linked to the source image
  • Region-level transcription with language and script metadata
  • Scene or environment context where relevant
  • Optional reading order or translation layers

Typical applications

OCR, scene-text detection, multilingual text recognition, visual translation and vision-language grounding.

Delivery note

Languages, scripts, scene types and annotation depth are defined around the target OCR or scene-text problem.

Delivery

The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.

Related datasets

Explore all collections