Learn how to run Kimi K2 Dynamic GGUFs Read our Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. 🌙 Kimi K2 Usage Guidelines You can now use the latest update of llama.cpp to run the model. For complete detailed instructions, see our guide: docs.unsloth.ai/basics/kimi k2 It is recommended to have at least 128GB unified RAM memory to run the small quants. With 16GB VRAM and 256 RAM, expect 5+ tokens/sec. For best results, use any 2 bit XL quant or above. Set the temperature to 0.6 recommended) to reduce repetition and incoherence. 📰 Tech Blog 📄 Paper Link (coming soon) 0. Changelog 2025.7.15 We have updated our tokenizer implementation. Now special tokens like [EOS] can be encoded to their token ids. We fixed a bug in the chat template that was breaking multi turn tool calls. 1. Model Introduction Kimi K2 is a state of the art mixture of experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticul…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy