Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 🤗 HuggingFace 🤖️ ModelScope 📔 Technical Report I. Introduction In this report, we share our efforts to build a compact yet powerful VLM, MiMo VL 7B. MiMo VL 7B comprises (1) a native resolution ViT encoder that preserves fine grained visual details, (2) an MLP projector for efficient cross modal alignment, and (3) our MiMo 7B language model, specifically optimized for complex reasoning tasks. The development of MiMo VL 7B involves two sequential training processes: (1) A four stage pre training phase, which includes projector warmup, vision language alignment, general multi modal pre training, and long context Supervised Fine Tuning (SFT). This phase yields the MiMo VL 7B SFT model. (2) A subsequent post training phase, where we introduce Mixed On policy Reinforcement Learning (MORL), a novel framework that seamlessly integrates diverse reward signals spanning perception accuracy, visual grounding precision, logical reasoning capabilities, and human/AI preferences. This…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy