OJ4OCRMT: A Large Multilingual Dataset for OCR MT Evaluation Check out the Paper: "OJ4OCRMT: A Large Multilingual Dataset for OCR MT Evaluation" Paul McNamee, Kevin Duh, Cameron Carpenter, Ron Colaianni, Nolan King, and Kenton Murray. Proceedings of Machine Translation Summit XX, Vol. 1: Research Track June 23 27, 2025, Geneva, Switzerland. The OJ4OCRMT dataset contains source PDF files, rendered images in three resolutions, and text files (both raw extractions, and sentence boundary split files). There are two partitions, 'dev' and 'test'. Each contains over 1,000 pages of content, with the PDFs, PNGs, and text files available in 23 EU languages. The dataset is designed to support evaluating systems for translation of document images between any pair of 23 European languages. The dev partition contains 1,656 pages from 2022. Of these 1,412 (85%) are deemed regular; 193 (12%) contain a table; and, 51 (3%) contain a 'figure'. The test partition contains 1,119 pages, all from 2023. 979 (87%) are regular; 98 (9%) contain a table; and, 42 (4%) contain a 'figure'. Note, all 2,772 pages have translations in all 23 languages. The languages are: Bulgarian, Croatian, Czech, Danish, Dutch, E…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy