GLM 5.1 Abliterated Dynamic IQ3 340 GGUF GGUF quantized version of helixdouble/GLM 5.1 Abliterated (base: zai org/GLM 5.1 FP8). 744B parameters MoE / 79 layers / 256 experts / 40B active / context length 202,752 ⚠️ CRITICAL: Known Issue — Garbage Output on CUDA (As of 2026 05 18) This model currently produces garbage output on all tested CUDA configurations. The build time "smoke test" only measured token generation speed (tok/s) — output quality was never verified . Symptom Greedy decoding ( temperature=0 ): always emits token 0 ( ! ) repeatedly: !!!!!!!!!!!!!!!!!!!! Sampling: random character garbage (e.g., H2%(@G&=6 9C4,' ) Affected Configurations (all tested, all failed) GPUs ctx size Key flags Result 6× RTX PRO 6000 Blackwell 4096 mla 3 muge merge qkv amb 512 ctk q8 0 ctv q8 0 flash attn on ❌ Garbage 4× RTX PRO 6000 Blackwell 30000 ctk q8 0 ctv q8 0 ❌ Garbage 4× RTX PRO 6000 Blackwell various cpu moe ❌ Garbage + slow 4× RTX PRO 6000 Blackwell various GGML CUDA FORCE MMQ=1 ❌ Garbage + slow 4× RTX PRO 6000 Blackwell various partial offload ngl 40 ❌ Garbage + slow 4× RTX PRO 6000 Blackwell various ctk f16 ctv f16 ❌ Garbage 6× RTX PRO 6000 Blackwell 4096 no mmap ❌ Garbage 6× RTX P…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy