OCR Synthetic Multilingual v1 Dataset Description Large scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state of the art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/AtharvImmverse/NemoTron OCR Synthetic Multilingual v1.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy