OCR Synthetic Multilingual v1 Dataset Description Large scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2 , a state of the art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection. This dataset is ready for commercial/non commercial use. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: April 15, 2026 License/Terms of Use: Dataset Governing Terms: Use of the dataset is governed by the Creative Commons Attribution 4.0 International License (CC BY 4.0). Intended Usage: This dataset is intended for machine learning researchers, AI engineers, and developers working on information retrieval with OCR. Dataset Characterization Data Collection Method [Hybrid: Human, Automated, Synthetic] Labeling Method [Not Applicable] Dataset Format Format — HDF5 Each .h5 file contains the following datasets (HDF5 terminology): Key Type Description images object (variable length bytes) JPEG encoded image bytes,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy