Molmo2 4B Molmo2 is a family of open vision language models developed by the Allen Institute for AI (Ai2) that support image, video and multi image understanding and grounding. Molmo2 models are trained on publicly available third party datasets as referenced in our technical report and Molmo2 data, a collection of datasets with highly curated image text and video text pairs. It has state of the art performance among multimodal models with a similar size. You can find all models in the Molmo2 family here. Learn more about the Molmo2 family in our announcement blog post. Molmo2 4B is based on Qwen3 4B Instruct and uses SigLIP 2 as vision backbone. It outperforms others in the class of open weight and data models on short videos, counting, and captioning, and is competitive on long videos. Ai2 is commited to open science. The Molmo2 datasets are available here. All other artifacts used in creating Molmo2 (training code, evaluations, intermediate checkpoints) will be made available at a later date, furthering our commitment to open source AI development and reproducibility. Quick links: 📂 All Models 📃 Paper 🎥 Blog with Videos Quick Start Setup Conda Environment General Video QA Poi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy