GLM 5.1 AWQ Base model: ZhipuAI/GLM 5.1 This repo quantizes the model using data free quantization tool. (no calibration dataset was involved) 【Dependencies / Installation】 As of 2026 04 20 , make sure your system has cuda12.8 or cuda13.0 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: vLLM Official Guide 【vLLM Startup Command】 Note: When launching with TP=8, include enable expert parallel ; otherwise the expert tensors wouldn’t be evenly sharded across GPU devices. 【Logs】 【Model Files】 File Size Last Updated 392GiB 2026 04 20 【Model Download】 【Overview】 GLM 5.1 👋 Join our WeChat or Discord community. 📖 Check out the GLM 5.1 blog and GLM 5 Technical report . 📍 Use GLM 5.1 API services on Z.ai API Platform. 🔜 GLM 5.1 will be available on chat.z.ai in the coming days. [ Paper ] [ GitHub ] Introduction GLM 5.1 is our next generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state of the art performance on SWE Bench Pro and leads GLM 5 by a wide margin on NL2Repo (repo generation) and Terminal Bench 2.0 (real world terminal tasks). But the most meaningful leap goes bey…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy