Introduction We are excited to introduce Ristretto , our newest Vision language model (VLM) that represents a significant step forward in the field. Ristretto features a capability to deploy dynamic image tokens, enables flexible adjustment of image token quantities based on task requirements while enhancing the projector architecture to support dynamic token configurations. This new model delivers improved performance and versatility compared to its predecessors through its refined architecture and advanced training approach. Environment Setup How to use? Evaluation Benchmark Qwen2.5 VL 3B InternVL2.5 4B Ristretto 3B : : : : : : : : MMBench TEST avg 76.8 78.2 80.1 MMStar 56.3 58.7 62.6 MMMU VAL 51.2 51.8 49.1 MathVista MINI test 61.2 60.8 67.9 HallucinationBench 46.6 46.6 50.2 AI2D 81.4 81.4 84.3 OCRBench 82.8 82.0 84.0 MMVet 60.0 61.5 61.8 Average 64.5 65.1 67.6 We use VLMEvalKit to evaluate Ristretto 3B. Other results are taken from OpenCompass License Agreement All of our open source models are licensed under the Apache 2.0 license. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy