Chandra OCR 2 Chandra 2 is a state of the art OCR model from Datalab that outputs markdown, HTML, and JSON. It is highly accurate at extracting text from images and PDFs, while preserving layout information. Try Chandra in the free playground, or use the hosted API for higher accuracy and speed. What's New in Chandra 2 85.9% olmocr bench score (sota), 77.8% multilingual bench score (12% improvement over Chandra 1) Significant improvements to math, tables, complex layouts Improved layout, especially on wider documents Significantly better image captioning 90+ language support with major accuracy gains Features Convert documents to markdown, HTML, or JSON with detailed layout information Excellent handwriting support Reconstructs forms accurately, including checkboxes Strong performance with tables, math, and complex layouts Extracts images and diagrams, with captions and structured data Support for 90+ languages Quickstart Usage With vLLM (recommended) With HuggingFace Transformers Benchmarks olmOCR Benchmark Model ArXiv Old Scans Math Tables Old Scans Headers and Footers Multi column Long tiny text Base Overall Source : : : : : : : : : : : : : : : : : : : : : Datalab API 90.4 90.2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy