Qwen3.6 27B Claude Opus Sonnet DistilledV2 MTP GGUF GGUF quantized release of the Claude Opus / Sonnet reasoning distillation on Qwen3.6 27B, with native MTP speculative decoding support in llama.cpp . Key numbers: Q4 K M + MTP2 → 114.78 tok/s generation, 80.33% draft acceptance, 64% faster than non MTP baseline. On the same machine, this release delivers 2x the visible answer content vs the original qwen3.6 27b while maintaining 4/4 correctness. Quick Download File Size Best for : : : Q4 K M (recommended) 15.66 GB Best overall balance Q6 K 20.89 GB Quality first Q2 K 10.12 GB Extreme compression Q8 0 27.05 GB High fidelity experiments Compared to Original qwen3.6 27b Same machine benchmark against the original (non quantized) qwen3.6 27b: GGUF side includes llama cli cold start — this is a conservative estimate. Original This release Average response time 10.93s 10.09s Correctness (4 prompts) 3/4 4/4 Visible answer chars 1336 2845 Hidden reasoning overhead 9002 chars minimal The original spends a large fraction of its token budget on hidden reasoning chains. This release converts that budget into visible answers, making it better suited for interactive local use. Compatibility Req…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy