🖥️ API / Platform      🗣️ Discord      🔗 GitHub      🤗 Demo Reasoning comes to OCR 🧠✨📄🤘 NuMarkdown 8B Thinking is the first reasoning OCR VLM. It is specifically trained to convert documents into clean Markdown files, well suited for RAG applications. It generates thinking tokens to figure out the layout of the document before generating the Markdown file. It is particularly good at understanding documents with weird layouts and complex tables. The number of thinking tokens can vary from 20% to 500% of the final answer, depending on the task difficulty. NuMarkdown 8B Thinking is a fine tune of Qwen 2.5 VL 7B on synthetic Doc → Reasoning → Markdown examples, followed by an RL phase (GRPO) with a layout centric reward. Try it out in the 🤗 space! Results NuMarkdown 8B Thinking is outperforming generic non reasoning models like GPT 4o and specialized OCR models like OCRFlux. It is competitive against large reasoning closed source models like Gemini 2.5. Arena ranking against popular alternatives (using trueskill 2 ranking system, with around 500 model anonymized votes): Rank Model μ σ μ − 3σ 🥇 1 gemini flash reasoning 2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy