ParseBench Quick links: [\[π Website\]](https://parsebench.ai) [\[π Paper\]](https://arxiv.org/abs/2604.08538) [\[π» Code\]](https://github.com/run llama/ParseBench) ParseBench is a benchmark for evaluating document parsing systems on real world enterprise documents, with the following characteristics: Multi dimensional evaluation. The benchmark is stratified into five capability dimensions β tables, charts, content faithfulness, semantic formatting, and visual grounding β each with task specific metrics designed to capture what agentic workflows depend on. Real world enterprise documents. The evaluation set contains ~2,000 human verified pages from over 1,200 publicly available documents spanning insurance, finance, government, and other domains, ranging from straightforward to adversarially hard. Dense test coverage. Over 169K test rules across the five dimensions, providing fine grained diagnostic power over precisely where a parser breaks down. Human verified annotations. All annotations are produced through a two pass pipeline: frontier VLM auto labeling followed by targeted human correction. Evaluation code suite. The benchmark ships with a full evaluation framework supportβ¦
Runs entirely in your browser via DuckDB-Wasm β this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy