Qwen3.6 14B A3B FableVibes GGUF GGUF quantizations of Qwen3.6 14B A3B FableVibes, a 14B MoE model fine tuned on reasoning traces from Claude Fable 5. Background This model started as Qwen3.6 35B A3B heretic and was pruned via REAP down to ~14B active parameters, removing over half its expert capacity. A single QLoRA pass was then orchestrated entirely by an autonomous AI agent ( Steve ), utilizing ~4,600 raw reasoning traces from Claude Fable 5 (Mythos class) to recover capabilities lost during pruning. Rather than focusing strictly on agentic orchestration, this model serves as a general purpose reasoning distill. The Fable CoT traces provide structured multi step reasoning patterns from a frontier class model, distilled into a footprint that can run on consumer hardware. The Fable traces are further supplemented by Claude Opus reasoning, Qwen tool calling data, and Evol Instruct Code. Available Formats Quant Size Notes F16 ~27GB Full precision reference Q8 0 ~15GB Near lossless Q6 K ~11.3GB Quality/size sweet spot Q5 K M ~9.8GB High quality BPW4.75 ~8.5GB Custom exl2 matched quantization array Q4 K M ~8.4GB Recommended for 8 12GB VRAM Q3 K M ~6.7GB Tight fits Q2 K ~5.3GB Maximum…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy