mlx community/GLM 5.2 DQ4plus q8 This model mlx community/GLM 5.2 DQ4plus q8 was converted to MLX format from zai org/GLM 5.2 using mlx lm version 0.31.3 (with PR 1410). This is created for people using a single Apple Mac Studio M3 Ultra with 512 GB. The 4 bit version of GLM 5.2 fits comfortably. But we can do better. Using research results, we aim to get better results from a slightly larger and smarter quantization. It should also not be so large that it leaves no memory for a useful context window. You can find more similar MLX model quants for Apple Mac Studio with 512 GB at https://huggingface.co/bibproj What is this DQ4plus q8? In the Arxiv paper Quantitative Analysis of Performance Drop in DeepSeek Model Quantization the authors write, We further propose DQ3 K M , a dynamic 3 bit quantization method that significantly outperforms traditional Q3 K M variant on various benchmarks, which is also comparable with 4 bit quantization ( Q4 K M ) approach in most tasks. and dynamic 3 bit quantization method ( DQ3 K M ) that outperforms the 3 bit quantization implementation in llama.cpp and achieves performance comparable to 4 bit quantization across multiple benchmarks. The resulting…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy