Model Overview Model Architecture: GLM 5.2 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.0.0 PyTorch: 2.9.0 Transformers: 5.8.1 Operating System(s): Linux Inference Engine: SGLang/vLLM Model Optimizer: AMD Quark (V0.11) Weight quantization: MOE only (shared experts quantized), OCP MXFP4, Static Activation quantization: MOE only, OCP MXFP4, Dynamic This model was built with GLM 5.2 model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from zai org/GLM 5.2 using AMD Quark. The weights and activations are quantized to MXFP4. Quantization scripts: Deployment Use with SGLang/vLLM This model can be deployed efficiently using the SGLang or vLLM backends. Evaluation The model was evaluated on GSM8K benchmarks. Accuracy Benchmark GLM 5.2 GLM 5.2 MXFP4(this model) Recovery GSM8K (flexible extract) 94.09 93.93 99.8% Reproduction The GSM8K results were obtained using the lm evaluation harness framework, based on the Docker image lmsysorg/sglang:v0.5.13.post1 rocm700 mi35x , with SGLang pre installed inside the image and lm eval compiled and installed from source. The Docker image rocm/vllm dev:nightly main 202606…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy