Infinity Doc2 5M 💻 Github 🤗 Infinity Parser2 Pro 🤗 Infinity Parser2 Flash 📄 Paper 🚀 Demo Infinity Doc2 5M is a training dataset for document parsing scenarios, with the following characteristics: Diverse document types : This dataset contains 5 million samples covering a wide range of document types, multiple layout types, and supports both Chinese and English languages. It encompasses academic papers, research reports and financial reports, newspapers, textbooks, exam papers, magazines, and more. Rich annotations : Includes detailed block level categories (titles, text paragraphs, tables, formulas, headers, footers, etc.), document element localization information, recognition results for each element region (text strings, table HTML, formula LaTeX, chemical SMILES, charts), and the overall reading order of the document. Diverse prompts : To address the lack of prompt diversity, we have constructed prompts with varied diversity. High data quality : Produced through manual filtering, intelligent annotation, and data synthesis. Manual annotation combined with expert quality inspection ensures high quality document image annotation data. Our corpus based data synthesis engine ca…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy