Ovis2.5 9B Introduction We are pleased to announce the release of Ovis2.5 , the successor to Ovis2, designed for native resolution visual perception and enhanced multimodal reasoning. It integrates a native resolution vision transformer (NaViT) that processes images at their original, variable resolutions, eliminating the need for fixed resolution tiling and preserving both fine details and global layout—crucial for visually dense content such as charts and diagrams. To strengthen reasoning, Ovis2.5 is trained not only on linear chain of thought (CoT) but also on reflective reasoning, including self checking and revision. This advanced capability is available at inference as an optional thinking mode , enabling users to trade latency for higher accuracy on complex inputs. Building on these advances, Ovis2.5 9B achieves an average score of 78.3 on the OpenCompass multimodal evaluation suite (SOTA among open source MLLMs under 40B parameters), while the lightweight Ovis2.5 2B scores 73.9, continuing the “small model, big performance” philosophy for resource constrained scenarios. Key Features Native Resolution Perception — NaViT vision encoder preserves fine details and global struct…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy