π€ HuggingFace π° Blog π¨ Xiaomi MiMo API Platform (Request Access) π¨οΈ Xiaomi MiMo Studio (Free Trial) Community WeChat Group Discord Telegram Reddit MiMo V2.5 Pro FP4 DFlash MiMo V2.5 Pro FP4 DFlash is the underlying model that powers MiMo V2.5 Pro UltraSpeed: An FP4 quantized backbone that applies MXFP4 quantization to the MoE experts while keeping the rest of the model at higher precision, shrinking model size and memory bandwidth pressure with near lossless quality. A BF16 DFlash drafter for block diffusion speculative decoding, which proposes a whole block of tokens per forward pass and lets the backbone verify them in one step. Together they cut both the per parameter bit width and the number of backbone forward passes, the two dominant costs of trillion parameter decoding. 1. Introduction At the trillion parameter (1T) scale, even 8 bit (FP8/INT8) inference carries severe memory footprint and memory bandwidth costs. Lowering the parameter bit width translates directly into faster decoding. We therefore adopt FP4 quantization and block diffusion speculative decoding. Key features of this release: Expert Onβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy