Infinity Parser 7B 💻 Github 📊 Dataset 📄 Paper 🚀 Demo Introduction We develop Infinity Parser, an end to end scanned document parsing model trained with reinforcement learning. By incorporating verifiable rewards based on layout and content, Infinity Parser maintains the original document's structure and content with high fidelity. Extensive evaluations on benchmarks in cluding OmniDocBench, olmOCR Bench, PubTabNet, and FinTabNet show that Infinity Parser consistently achieves state of the art performance across a broad range of document types, languages, and structural complexities, substantially outperforming both specialized document parsing systems and general purpose vision language models while preserving the model’s general multimodal understanding capability. Key Features LayoutRL Framework: a reinforcement learning framework that explicitly trains models to be layout aware through verifiable multi aspect rewards combining edit distance, paragraph accuracy, and reading order preservation. Infinity Doc 400K Dataset: a large scale dataset of 400K scanned documents that integrates high quality synthetic data with diverse real world samples, featuring rich layout variations…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy