Granite 4.0 3B Vision Model Summary: Granite 4.0 3B Vision is a vision language model (VLM) designed for enterprise grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with: Chart extraction: Converting charts into structured, machine readable formats (Chart2CSV, Chart2Summary, and Chart2Code) Table extraction: Accurately extracting tables with complex layouts from document images to JSON, HTML, or OTSL Semantic Key Value Pair (KVP) extraction: Extracting values based on key names and descriptions across diverse document layouts The model is delivered as a LoRA adapter on top of Granite 4.0 Micro, with a 3.5B base LLM and 0.5B LoRA adapters. This enables a single deployment to support both multimodal document understanding and text only workloads — the base model handles text only requests without loading the adapter. See Model Architecture for details. The methodology and data (ChartNet) used for this model are described in the paper ChartNet: A Million Scale, High Quality Multimodal Dataset for Robust Chart Understanding. While our focus is on specialized document extraction tasks, the current model preserves…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy