IMPORTANT: You must use the docker image below, since it contains many custom kernels written for this model specifically Updated 5/4/26 New calibration data, new docker image. Full 2x RTX 6K support Model Description MiMo V2.5 NVFP4 is an NVFP4 quantized version of XiaomiMiMo/MiMo V2.5. This is a multi modal model, supporting text, images, audio and video. This quantization carefully preserves those capabilities. What's quantized Only the non shared MoE expert MLP projections are quantized to NVFP4. Attention weights are left in BF16, in addition to the dense MLPs (layers 0 3) and the shared experts. Since the MoE expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings. Calibration uses natural top k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a much larger number of samples than typical to ensure broad expert coverage through natural routing alone. Calibration dataset Six calibration passes were run: 1. Coding — Agentic coding samples (tool calling, multi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy