GLM 5.2 REAP50 Q3 K M GGUF A GGUF build of GLM 5.2, REAP expert pruned (50%) and quantized to Q3 K M (~169 GB) — sized to run on 2× 96 GB GPUs (e.g. RTX PRO 6000), ~192 GB VRAM , with room for context. What this is Base: zai org/GLM 5.2 ( glm moe dsa , ~753B MoE). REAP 50 : the 128 most salient experts per layer kept (of 256) via Cerebras REAP saliency ( gate × ‖expert output‖ ), MTP layer dropped → ~394B params. Quantized to Q3 K M , split into 5 shards (~45 GB each). Runs as full MLA attention (the DSA lightning indexer is not used at inference — same simplification as the upstream conversion). ⚠️ Requires a patched llama.cpp (for now) Stock llama.cpp can't load any GLM 5.2 GGUF yet: its GLM DSA loader requires the DSA indexer tensors on every layer, but GLM 5.2 only ships them on a subset ("full") of layers → missing tensor 'blk.N.indexer.k norm.weight' . The indexer is loaded but unused (the graph is DeepSeek V2 MLA), so the fix is simply to make those tensors optional. Apply the included llama.cpp glm dsa indexer optional.patch ( src/models/glm dsa.cpp ) and rebuild, or wait for the upstream GLM DSA runtime PR. After patching it loads and runs normally. Quality caveat This is…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy