GLM 5.2 Int4 Int8Mix Base model: zai org/GLM 5.2 This repo quantizes the model using a data free quantization tool. (no calibration dataset was involved) This release is prepared for vLLM with compressed tensors weight only A16 inference. It does not claim compatibility with SGLang or other runtimes. The default reasoning effort is changed to medium high to reduce thinking token cost. You can override it per request with "chat template kwargs": {"reasoning effort": "max"} , or edit chat template.jinja to set a different default. 【Quantization Policy】 Scope Format model.layers.0 BF16 model.layers.1 model.layers.2 W8A16, group size 128 model.layers.3 model.layers.77 ordinary linear weights W8A16, group size 128 model.layers.3 model.layers.77 MoE expert weights W4A16, group size 128 model.layers.78 MTP block W8A16, channelwise mlp.gate. FP32 Attention indexer, norms, embeddings, and special heads BF16 Accuracy reference, using the SGLang GLM 5.2 FP8 H200 / default / low latency / single node AIME25 recipe: Model Runtime Quantization Reasoning effort AIME25 pass@1 ZhipuAI/GLM 5.2 FP8 SGLang FP8 max 87.7% tclf90/GLM 5.2 Int4 Int8Mix vLLM Int4 Int8Mix, W4A16/W8A16 max 92.92% tclf90/GLM 5…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy