Chandra Chandra is an OCR model that outputs markdown, HTML, and JSON. It is highly accurate at extracting text from images and PDFs, while preserving layout information. You can try Chandra in the free playground here, or at a hosted API here. Features Convert documents to markdown, html, or json with detailed layout information Good handwriting support Reconstructs forms accurately, including checkboxes Good support for tables, math, and complex layouts Extracts images and diagrams, with captions and structured data Support for 40+ languages Quickstart The easiest way to start is with the CLI tools: Benchmarks We used the olmocr benchmark, which seems to be the most reliable current OCR benchmark in our testing. Model ArXiv Old Scans Math Tables Old Scans Headers and Footers Multi column Long tiny text Base Overall Source : : : : : : : : : : : : : : : : : : : : : Datalab Chandra v0.1.0 82.2 80.3 88.0 50.4 90.8 81.2 92.3 99.9 83.1 ± 0.9 Own benchmarks Datalab Marker v1.10.0 83.8 69.7 74.8 32.3 86.6 79.4 85.7 99.6 76.5 ± 1.0 Own benchmarks Mistral OCR API 77.2 67.5 60.6 29.3 93.6 71.3 77.1 99.4 72.0 ± 1.1 olmocr repo Deepseek OCR 75.2 72.3 79.7 33.3 96.1 66.7 80.1 99.7 75.4 ± 1.0 O…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy