[!NOTE] Includes Unsloth chat template fixes ! For llama.cpp , use jinja Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. 𓌳 REAP 𓌳 the Experts: Why Pruning Prevails for One Shot MoE Compression GLM 4.7 Flash REAP 23B A3B ✨ Highlights Introducing GLM 4.7 Flash REAP 23B A3B , a memory efficient compressed variant of GLM 4.7 Flash that maintains near identical performance while being 25% lighter . This model was created using REAP (Router weighted Expert Activation Pruning) , a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts. Key features include: Near Lossless Performance : Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 30B model 25% Memory Reduction : Compressed from 30B to 23B parameters, significantly lowering deployment costs and memory requirements Preserved Capabilities : Retains all core functionalities including code generation, agentic workflows, repository scale understanding, and function calling Drop in Compatibility : Works with vanilla vLLM no source modifications or c…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy