OmniDocBench English 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real world scenarios, with the following characteristics: Diverse Document Types : The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations : Contains localization for 28 block level categories (text paragraphs, titles, tables, formulas, headers/footers, etc.) and 4 span level categories (text lines, inline formulas, superscripts/subscripts, etc.), plus recognition results for each region (text, LaTeX for formulas, LaTeX and HTML for tables). OmniDocBench also provides reading order annotations for layout elements. Page and block level attribute labels include 5 page attribute categories, 3 text related attributes and 6 table related attributes. High Annotation Quality : Through manual screening, intelligent annotation, manual annotation, full expert quality inspection and large model quality inspection, the data quality is relatively high. Evaluation Co…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy