Product and Packaging Image Sets
Product imagery across views, packaging surfaces and visual conditions for recognition, retrieval and packaging understanding.
View datasetScene-text models need more than an image containing words. They need to know where the text appears, what it says and which language or script it uses. Real-World Text and Signage Photos make those relationships explicit at the region level.
Text geometry can be connected to transcription and scene context, with reading order or translation added where the task benefits from them. This creates a direct image-to-region-to-text structure for models working outside the clean conditions of scanned documents.
OCR, scene-text detection, multilingual text recognition, visual translation and vision-language grounding.
Languages, scripts, scene types and annotation depth are defined around the target OCR or scene-text problem.
The exact package depends on the collection scope. Where relevant, delivery can include task-specific records, manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions.