GLM 5.2 — W4A16 (INT4) + BF16 MTP An INT4 weight only (W4A16) quantization of GLM 5.2 that preserves the BF16 multi token prediction (MTP) layer for speculative decoding. Quantized from zai org/GLM 5.2 with llm compressor (GPTQ). Built for Hopper (H200), validated on Blackwell (8×RTX PRO 6000, SM120). Matches FP8 quality on half the GPUs (4×H200 vs 8) and is the fastest GLM 5.2 quant for interactive/agentic serving on Hopper in a matched, MTP on head to head — with the lowest time to first token by a wide margin. A complete, quality validated RTX PRO 6000 recipe is in the Serving on Blackwell section below. Why this model Half the footprint, FP8 quality. ~405 GB of weights (down from ~1.49 TB BF16) serve one replica on 4×H200 instead of 8 — freeing half the fleet, or two replicas per node — and eval matches the FP8 baseline within noise across reasoning, instruction following, long context, and agentic coding. Fastest interactive serving among GLM 5.2 quants on Hopper. In a matched benchmark (every model with MTP on , same box, same vLLM, same harness): +8% vs nvidia NVFP4 and +33% vs zai FP8 at concurrency 1 , with TTFT of 215 ms vs 632/1258 ms . Honest trade off. MTP's draft/veri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy