Molmo 7B D Molmo is a family of open vision language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly curated image text pairs. It has state of the art performance among multimodal models with a similar size while being fully open source. You can find all models in the Molmo family here. Learn more about the Molmo family in our announcement blog post or the paper. Molmo 7B D is based on Qwen2 7B and uses OpenAI CLIP as vision backbone. It performs comfortably between GPT 4V and GPT 4o on both academic benchmarks and human evaluation. It powers the Molmo demo at molmo.allenai.org . This checkpoint is a preview of the Molmo release. All artifacts used in creating Molmo (PixMo dataset, training code, evaluations, intermediate checkpoints) will be made available at a later date, furthering our commitment to open source AI development and reproducibility. Sign up here to be the first to know when artifacts are released. Quick links: 💬 Demo 📂 All Models 📃 Paper 🎥 Blog with Videos Quick Start To run Molmo, first install dependencies: Then, follow these steps: To make inference more efficient, run with autocast: We did mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy