m51Lab MiniMax M2.7 REAP 139B A10B NVFP4 GGUF NVFP4 mixed precision GGUF quantization of the REAP 40% pruned MiniMax M2.7 model, optimised for NVIDIA Blackwell GPUs. The NVFP4 tensor types use native Blackwell FP4 hardware acceleration for maximum inference speed, while smaller tensors (norms, biases, router, embeddings) are preserved in Q4 K for accuracy. About the Source Model m51Lab MiniMax M2.7 REAP 139B A10B is a REAP 40% pruned variant of MiniMax M2.7, reducing total parameters from 229B to 139B while preserving the 10B active parameters per token. REAP (Router weighted Expert Activation Pruning) prunes 40% of experts per MoE block (256 → 154) based on router gated activation patterns. Base model: MiniMaxAI/MiniMax M2.7 (229B MoE, 62 layers) Pruning method: REAP (Lasby et al., 2025, arXiv:2510.13999) Pruning rate: 40% of experts per MoE block Active parameters: ~10B per token Architecture: Sparse MoE decoder only Transformer Hidden dim: 3072 Attention heads: 48 query / 8 KV (GQA) Experts: 154 per block (top 8 activated) Expert routing: Sigmoid gating with learnable bias Context length: 196,608 tokens Vocab: 200,064 (GPT 2 tokenizer) Evaluation (from source) HumanEval pass@1 (…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy