MiniMax M2.7 REAP 172B A10B NVFP4 GB10 A fully quantized compressed tensors checkpoint of catplusplus/MiniMax M2.7 REAP 172B A10B NVFP4 , reshaped to match the layout of saricles/MiniMax M2.5 REAP 172B A10B NVFP4 GB10 and tuned for the NVIDIA GB10 Grace Blackwell (SM 12.1) Spark node, while remaining loadable on upstream vLLM and SGLang without patches. What changed vs. the source catplusplus checkpoint Component Before (ModelOpt) After (compressed tensors) Effect MoE experts ( w1/w2/w3 ) packed FP4 with weight / weight scale / weight scale 2 / input scale renamed + scalar inversion: weight packed / weight scale / weight global scale / input global scale format parity with saricles; zero numeric change (round trip < 1e 8) Attention Q/K/V/O float32 (~9.8 GB unquantized) NVFP4 W4A4 compressed tensors, group size=16 ~8.5 GB saved; bandwidth bound on GB10 → decode throughput up lm head bfloat16 FP8 per tensor, compressed tensors float quantized ~0.6 GB saved; upstream compatible (no runtime env var hooks) k proj.k scale / v proj.v scale present (static FP8 KV) dropped KV cache quant handled by runtime (sglang / vLLM) Embeddings, norms, gate routers, bias bfloat16 bfloat16 (unchanged) r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy