Read our How to Run GLM 4.7 Flash Guide! Jan 21 update: llama.cpp fixed a bug that caused looping and poor outputs. We updated the GGUFs please re download the model for much better outputs. Repeat penalty: Disable it, or set repeat penalty 1.0 You can now use Z.ai's recommended parameters and get great results: For general use case: temp 1.0 top p 0.95 For tool calling: temp 0.7 top p 1.0 If using llama.cpp, set min p 0.01 as llama.cpp's default is 0.05 You can also fine tune GLM 4.7 Flash with Unsloth via our GLM free notebook. Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. GLM 4.7 Flash π Join our Discord community. π Check out the GLM 4.7 technical blog , technical report(GLM 4.5) . π Use GLM 4.7 Flash API services on Z.ai API Platform. π One click to GLM 4.7 . Introduction GLM 4.7 Flash is a 30B A3B MoE model. As the strongest model in the 30B class, GLM 4.7 Flash offers a new option for lightweight deployment that balances performance and efficiency. Performances on Benchmarks Benchmark GLM 4.7 Flash Qwen3 30B A3B Thinking 2507 GPT OSS 20B AIME 25 91.6 85.0 91.7 GPQA 75.2 73.4 71.5 LCB v6 64.0 66.0 61.0 HLE 14.4 9.8 10.9 SWE bench Verifβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy