GLM 4.7 AWQ Base model: zai org/GLM 4.7 【Dependencies / Installation】 As of 2025 12 24 , make sure your system has cuda12.8 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: 【vLLM Startup Command】 Note: When launching with TP=8, include enable expert parallel ; otherwise the expert tensors couldn’t be evenly sharded across GPU devices. 【Logs】 【Model Files】 File Size Last Updated 181 GiB 2025 12 24 【Model Download】 【Overview】 GLM 4.7 👋 Join our Discord community. 📖 Check out the GLM 4.7 technical blog , technical report(GLM 4.5) . 📍 Use GLM 4.7 API services on Z.ai API Platform. 👉 One click to GLM 4.7 . Introduction GLM 4.7 , your new coding partner, is coming with the following features: Core Coding : GLM 4.7 brings clear gains, compared to its predecessor GLM 4.6, in multilingual agentic coding and terminal based tasks, including (73.8%, +5.8%) on SWE bench, (66.7%, +12.9%) on SWE bench Multilingual, and (41%, +16.5%) on Terminal Bench 2.0. GLM 4.7 also supports thinking before acting, with significant improvements on complex tasks in mainstream agent frameworks such as Claude Code, Kilo Code, Cline, and Roo Code. Vibe Coding : GLM 4.7 takes a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy