Mellum2 12B A2.5B Thinking GGUF (Q4 K M) Quantized GGUF version of JetBrains/Mellum2 12B A2.5B Thinking for use with llama.cpp. ⚠️ Requires llama.cpp with Mellum architecture support. This is not yet in mainline llama.cpp — use PR 23966 or a build that includes it. Model Details Property Value Base model JetBrains/Mellum2 12B A2.5B Thinking Architecture Mellum (MoE) Total parameters 12B Active parameters 2.5B Experts 64 (8 per token) Context length 128K (sliding window) Quantization Q4 K M File size ~7.5 GB Usage For thinking mode, use reasoning budget to cap reasoning tokens: Low VRAM (6GB GPU) For GPUs with limited VRAM, offload MoE experts to CPU: Tested on GTX 1060 6GB: ~18 t/s with ncmoe=20, ~22 t/s with ncmoe=10. Acknowledgements JetBrains for the original Mellum2 model llama.cpp for the inference engine
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy