🤗 HuggingFace 🖥️ Demo 📄 Technical Report 🐈 GitHub Figure 1: Performance comparison on the OmniDocBench v1.5 benchmark. FireRed OCR achieves state of the art performance among end to end solutions, ranking first with a score above 92%. 🔥 FireRed OCR FireRed OCR is a systematic framework designed to specialize general Large Vision Language Models (LVLMs) into high performance, pixel precise structural document parsing experts. General VLMs frequently suffer from "Structural Hallucination" (e.g., disordered rows, invented formulas) when processing complex documents. FireRed OCR addresses this by shifting the paradigm from "impressionist" text generation to "structural engineering," achieving State of the Art (SOTA) results on authoritative benchmarks like OmniDocBench v1.5. ✨ Key Features SOTA Performance : Achieves 92.94% overall score on OmniDocBench v1.5, significantly outperforming DeepSeek OCR 2, OCRVerse, and massive general VLMs (e.g., Gemini 3.0 Pro,Qwen3 VL 235B). Structural Integrity : Utilizing Format Constrained GRPO (Group Relative Policy Optimization), the model enforces strict syntactic validity, eliminating common errors like unclosed tables or invalid LaTeX formu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy