GLM 5.2 NVFP4 NVFP4 (4 bit) quantization of zai org/GLM 5.2 , produced with NVIDIA TensorRT Model Optimizer 0.44.0. The MoE expert FFNs (routed + shared) are quantized to NVFP4; attention (MLA + the DeepSeek style DSA lightning indexer), the router, and the LM head are kept in BF16. This shrinks the checkpoint from 1.5 TB → 410 GB (~3.7×) while retaining GSM8K accuracy within ~2 points of BF16. GLM 5.2 is a glm moe dsa model: DeepSeek V3.2 style MLA attention + DSA sparse attention indexer , with a 256 routed expert + 1 shared expert MoE (8 experts/token), 78 layers, hidden 6144, vocab 154880. Evaluation All benchmarks were served via SGLang and scored with lm evaluation harness on the same hardware and harness for both NVFP4 and BF16 (generative / chain of thought where applicable; max gen toks raised to fit the reasoning chains — lm eval's default 256 truncates them and tanks the scores). Benchmark GLM 5.2 NVFP4 (410 GB) GLM 5.2 BF16 (1507 GB) Δ GPQA Diamond (CoT, flexible) 69.70 69.70 0.00 MATH 500 (minerva) 86.80 86.60 +0.20 MMLU Pro (generative, 50/subject) 81.14 82.43 −1.29 HumanEval (pass@1, instruct) 94.51 95.73 −1.22 GSM8K (5 shot, flexible) 92.72 94.92 −2.20 NVFP4 holds u…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy