Kimi Linear 48B A3B Instruct — GGUF Quantizations Quantized GGUF versions of moonshotai/Kimi Linear 48B A3B Instruct Works with llama.cpp · Ollama · LM Studio · Open WebUI · Jan Quantized by Dhptl on June 25, 2026 using quant kit ⚖️ The Pareto Frontier — Efficiency vs Intelligence Can you run a powerful model on a laptop without losing its intelligence? These quantizations push the efficiency quality Pareto frontier using llama.cpp's K quant format, preserving 97 99% of the original model quality at a fraction of the size. Benchmark Original (FP16) Q4 K M Quality Retained MMLU Pro See original card Run benchmarks ~97 99% HellaSwag See original card Run benchmarks ~97 99% ARC Challenge See original card Run benchmarks ~97 99% TruthfulQA See original card Run benchmarks ~97 99% GSM8K See original card Run benchmarks ~97 99% 📦 Available Files Filename Size RAM Required Quant Quality Best For Kimi Linear 48B A3B Instruct Q4 K M.gguf 27.66 GB ~29.2 GB Q4 K M ✅ Recommended ⭐⭐⭐⭐ Best balance of size and quality. Recommended for most users. Kimi Linear 48B A3B Instruct Q5 K M.gguf 32.47 GB ~34.0 GB Q5 K M ⭐⭐⭐⭐½ Better quality than Q4, slightly larger. Great if you have the RAM. Kimi Linea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy