Qianfan OCR A Unified End to End Model for Document Intelligence π€ Demo π Technical Report π₯οΈ Qianfan Platform π» GitHub π§© Skill Introduction Qianfan OCR is a 4B parameter end to end document intelligence model developed by the Baidu Qianfan Team. It unifies document parsing, layout analysis, and document understanding within a single vision language architecture. Unlike traditional multi stage OCR pipelines that chain separate layout detection, text recognition, and language comprehension modules, Qianfan OCR performs direct image to Markdown conversion and supports a broad range of prompt driven tasks β from structured document parsing and table extraction to chart understanding, document question answering, and key information extraction β all within one model. Key Highlights π 1 End to End Model on OmniDocBench v1.5 β Achieves 93.12 overall score, surpassing DeepSeek OCR v2 (91.09), Gemini 3 Pro (90.33), and all other end to end models π 1 End to End Model on OlmOCR Bench β Scores 79.8 π 1 on Key Information Extraction β Overall mean score of 87.9 across five public KIE benchmarks, surpassing Gemini 3.1 Pro, Gemini 3 Pro, Seed 2.0, and Qwen3 VL 235B A22B π§ Layout as Thoβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy