mlx community/Kimi K2.6 mlx DQ3 K M q8 This model mlx community/Kimi K2.6 mlx DQ3 K M q8 was converted to MLX format from moonshotai/Kimi K2.6 using mlx lm version 0.31.2 . After the success of the first Kimi "DQ3 K M" model and the K2.5, this is a new update for Kimi K2.6! This is created for people using a single Apple Mac Studio M3 Ultra with 512 GB. The 4 bit version of Kimi K2 does not fit. Using research results, we aim to get 4 bit performance from a slightly smaller and smarter quantization. It should also not be so large that it leaves no memory for a useful context window. You can find more similar MLX model quants for Apple Mac Studio with 512 GB at https://huggingface.co/bibproj What is this DQ3 K M? In the Arxiv paper Quantitative Analysis of Performance Drop in DeepSeek Model Quantization the authors write, We further propose DQ3 K M , a dynamic 3 bit quantization method that significantly outperforms traditional Q3 K M variant on various benchmarks, which is also comparable with 4 bit quantization ( Q4 K M ) approach in most tasks. and dynamic 3 bit quantization method ( DQ3 K M ) that outperforms the 3 bit quantization implementation in llama.cpp and achieves perfor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy