🪷 NayanaOCR Corpus 2025 A 1M page, 22 language fully parallel synthetic OCR + VQA corpus for document centric vision language models — every page rendered in every language. NayanaOCR Corpus 2025 is one of the largest open source multilingual, multi task document datasets for training and evaluating OCR , layout detection , and visual question answering (VQA) in low resource and underrepresented languages. The headline property: it's a true parallel corpus. The same ~45,700 source pages are re rendered into all 22 languages with consistent region IDs and bounding boxes, joined by a shared image id.txt . That means every page exists in all 22 scripts — Latin, Devanagari, Dravidian, CJK, Arabic, Thai — giving you free cross lingual pairs, script invariant supervision, and clean cross script ablations out of the box (see § Why a Parallel Corpus? below). In total: ~1,006,170 richly annotated document page renderings across 22 languages , with per region bounding boxes, English ground truth, in language translations, and synthetic VQA pairs (descriptive + multiple choice). This dataset is part of the Nayana initiative — a 2025 recipient of the Meta Llama Impact Grant . We're grateful t…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy