Moondream is a small vision language model designed to run efficiently everywhere. Website / Demo / GitHub This repository contains the 2025 04 14 4 bit release of Moondream. On an Nvidia RTX 3090, it uses 2,450 MB of VRAM and runs at a speed of 184 tokens/second. We used quantization aware training techniques to build this version of the model, allowing us to achieve a 42% reduction in memory usage with only an 0.6% drop in accuracy. There's more information about this version of the model in our release blog post. Other revisions, as well as release history, can be found here. Usage Make sure to install the requirements:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy