[!NOTE] Includes Unsloth chat template fixes ! For llama.cpp , use jinja Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. 🤗 HuggingFace 📔 Technical Report 📰 Blog Play around! 🗨️ Xiaomi MiMo Studio 🎨 Xiaomi MiMo API Platform MiMo V2 Flash MiMo V2 Flash is a Mixture of Experts (MoE) language model with 309B total parameters and 15B active parameters . Designed for high speed reasoning and agentic workflows, it utilizes a novel hybrid attention architecture and Multi Token Prediction (MTP) to achieve state of the art performance while significantly reducing inference costs. 1. Introduction MiMo V2 Flash creates a new balance between long context modeling capability and inference efficiency. Key features include: Hybrid Attention Architecture : Interleaves Sliding Window Attention (SWA) and Global Attention (GA) with a 5:1 ratio and an aggressive 128 token window. This reduces KV cache storage by nearly 6x while maintaining long context performance via learnable attention sink bias . Multi Token Prediction (MTP) : Equipped with a lightweight MTP module (0.33B params/block) using dense FFNs. This triples output sp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy