[!NOTE] Includes Unsloth chat template fixes ! For llama.cpp , use jinja Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. 𓌳 REAP 𓌳 the Experts: Why Pruning Prevails for One Shot MoE Compression GLM 4.7 REAP 218B A32B ✨ Highlights Introducing GLM 4.7 REAP 218B A32B , a memory efficient compressed variant of GLM 4.7 that maintains near identical performance while being 40% lighter . This model was created using REAP (Router weighted Expert Activation Pruning) , a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts. Key features include: Near Lossless Performance : Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 355B model 40% Memory Reduction : Compressed from 355B to 218B parameters, significantly lowering deployment costs and memory requirements Preserved Capabilities : Retains all core functionalities including code generation, agentic workflows, repository scale understanding, and function calling Drop in Compatibility : Works with vanilla vLLM no source modifications or custom patch…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy