granite docling 258m Granite Docling is a multimodal Image Text to Text model engineered for efficient document conversion. It preserves the core features of Docling while maintaining seamless integration with DoclingDocuments to ensure full compatibility. Model Summary : Granite Docling 258M builds upon the Idefics3 architecture, but introduces two key modifications: it replaces the vision encoder with siglip2 base patch16 512 and substitutes the language model with a Granite 165M LLM. Try out our Granite Docling 258 demo today. Developed by : IBM Research Model type : Multi modal model (image+text to text) Language(s) : English (NLP) License : Apache 2.0 Release Date : September 17, 2025 Granite docling 258M is fully integrated into the Docling pipelines, carrying over existing features while introducing a number of powerful new features, including: 🔢 Enhanced Equation Recognition: More accurate detection and formatting of mathematical formulas 🧩 Flexible Inference Modes: Choose between full page inference, bbox guided region inference 🧘 Improved Stability: Tends to avoid infinite loops more effectively 🧮 Enhanced Inline Equations: Better inline math recognition 🧾 Document E…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy