MiniMax M2.7 REAP 172B A10B NVFP4 GB10 NVFP4 GB10 quantization of a 25% REAP pruned MiniMax M2.7, targeted at NVIDIA DGX Spark (GB10) and Blackwell family hardware. Both the REAP pruning AND the NVFP4 calibration use a 6 dataset agentic mix for coherent preservation of tool use, code generation, math, and software engineering agent capabilities. 98.9 GB on disk — fits in a single 128 GB DGX Spark. Model Details Base Model (BF16) saricles/MiniMax M2.7 REAP 172B A10B BF16 Original Base MiniMaxAI/MiniMax M2.7 Architecture MiniMaxM2ForCausalLM (MoE, 192 experts, top K=8) Total Parameters 172B (REAP pruned from 230B) Active Parameters ~10B per token Hidden Layers 62 Quantization NVFP4 (4 bit floating point) with GB10 tuned ignore list Format compressed tensors (safetensors) Size on Disk 98.9 GB Deployment 1× DGX Spark (fits in a single 128 GB Spark) License Other (inherited from MiniMaxAI/MiniMax M2.7) Lineage 1. Base: MiniMaxAI/MiniMax M2.7 (230B, 256 experts, FP8 native) 2. Dequantize: FP8 → BF16 3. REAP prune: BF16 → 75% keep (192/256 experts), agentic 6 dataset calibration → saricles/MiniMax M2.7 REAP 172B A10B BF16 4. NVFP4 quantize (this model): REAP BF16 + same 6 dataset agentic…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy