MiMo V2.5 — AWQ W4A16 (int4), A100 ready 4 bit ( AWQ W4A16, routed experts only ) quantization of MiMo V2.5 , packaged to serve on NVIDIA A100 (SM80) under stock vLLM 0.21.0 — text + vision, at TP 4 or TP 8. The base model ships fp8 (Hopper native) and does not run on A100. This repo is the A100 path: the int4 weights plus the two small vLLM model code patches that make MiMo's Hopper only attention run on Ampere — without changing the math . bf16 → int4 : ≈581 GB → 169 GB Context: up to 1,048,576 tokens · text + vision · tool calling + reasoning Use this model (vLLM) MiMo's default vLLM path hard selects a FlashAttention 3 backend (SM90+ only). Bind mount the two patch files over the image copies (details in PATCHES.md ): TP 4 works too — set tensor parallel size 4 . (The patches are correct at both; see the QKV note in PATCHES.md .) Sampling: temperature 1.0, top p 0.95 (the model's shipped generation config.json ; thinking mode on). max model len can be raised toward the native 1,048,576 as VRAM allows. Files file what model 0000{1..4} of 00004.safetensors int4 weights (W4A16) config.json , recipe.yaml quant config + the full quantization recipe modeling mimo v2.py , configuratio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy