Qwen 3.6 35B A3B — Cerebellum GGUF Sensitivity guided mixed precision quantization of Qwen/Qwen3.6 35B A3B. Cerebellum measures which weight groups survive extreme compression and which don't, then writes a single GGUF with per tensor precision assignments — a standard GGUF that runs on stock llama.cpp , no fork. Variant File Size BPW Best for 14 GB (recommended) Qwen3.6 35B A3B Cerebellum 14GB.gguf 14.0 GB 3.34 best coding, 160K+ context v3 (smallest) Qwen3.6 35B A3B Cerebellum v3 Q3 K M.gguf 11 GB 2.76 tightest VRAM, vision Evaluations Coding — upstream EvalPlus ( evalplus.codegen against llama server , greedy / temp 0, n=164), same protocol across the size ladder: build size HumanEval HumanEval+ : : : : micro 11.96 GB 90.9 87.2 14 GB (recommended) 14.0 GB 93.3 90.2 uniform Q3 K M 16.0 GB 91.5 89.0 Base 17.3 GB 92.7 89.0 Long context: needle recall passes to 90K+ ( verify stress ). Throughput: ~168 tok/s decode (3B active MoE); fits 160K+ context at ~19 GB on a 24 GB card. Per question artifacts in benchmark results/14gb/ . Why the 14 GB over v3 v3 (11 GB) is the tightest VRAM build. The 14 GB spends ~3 GB more to promote the routed ffn down exps to Q4 K — the group the ablation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy