RedHatAI/GLM 5.2 NVFP4 FP8 This is a quantized version of zai org/GLM 5.2 with MoE layers quantized to NVFP4 and attention layers quantized to FP8 block Usage This model is intended for deployment with vLLM and requires the following fix: https://github.com/vllm project/vllm/pull/47780. You can serve the model using Creation Process This model was created using LLM Compressor. The example script can be found in examples/quantizing moe/glm5 example.py [[Example] GLM5.2 Example](https://github.com/vllm project/llm compressor/pull/2869). Quantizing the model with data parallelism and 6xA100 takes about 3 hours. LLM Compressor Creation Script Evaluation Benchmark zai org/GLM 5.2 RedHatAI/GLM 5.2 NVFP4 FP8 GPQA Diamond 91.2 89.1
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy