MiniCPM5 1B Agentic v8 Created by GLM 5.2. Model 8 of 8 in the agentic post training series. Variant : 4 way Soup (best overall) Evaluation Metric Score Real World Tasks 33.3% (4/12) Unique Tasks Solved 8/12 Consistent (5/5) rw http server Often (4/5) rw venv setup GGUFs available: f16, q8 0, q5 k m, q4 k m, q3 k m, q2 k Quantization Recommendations This is a 1B model — heavier quantization degrades output quality significantly. Quant Quality Size Recommendation f16 Full ~2.1GB Best quality q8 0 Excellent ~1.1GB Recommended — near identical to f16 q5 k m Good ~0.8GB Reasoning OK, response may degrade on longer outputs q4 k m Fair ~0.7GB Reasoning OK, response degrades into repetition q3 k m Poor ~0.6GB Not recommended q2 k Poor ~0.5GB Not recommended For production use, prefer q8 0 or f16. The model uses reasoning tokens; lower quantizations break the transition from reasoning to response. Chat Template The GGUF chat template defaults to enable thinking=true , so the model will always produce reasoning followed by response. If your inference engine supports enable thinking=false , you can skip reasoning for faster responses.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy