Qwen3.6 35B A3B DSV4Pro Thinking Distill Lynn Agent edge runtime This 35B A3B distillation is the high end local orchestrator for Lynn Agent — the sparse (3B active) sibling of the 27B Dense distill. On 32 GB+ VRAM / unified memory machines Lynn can run it as the task orchestrator (decompose → delegate → verify); 24 GB machines use the 27B path instead. Download Lynn Agent : GitHub Releases v0.85.6 GGUF (edge) : Hugging Face 35B GGUF / ModelScope 35B GGUF FP8 (concurrent serving) : Hugging Face 35B FP8 Recommended quant : Q5 K M imatrix + native MTP (single stream), or FP8 for concurrent serving. 🇬🇧 English · 🇨🇳 中文 ⬇️ 🇬🇧 English On Qwen3.6 35B A3B ( MoE, 3B active ), we use LoRA to distill the way DeepSeek V4 Pro reasons (with thinking on) plus its agentic behavior — purpose built as a fast task orchestrator (decompose → delegate → verify) for Lynn Agent. This is the MoE counterpart of the 27B Dense sister model : same R6000 GPU, same teacher, same recipe, on a sparse architecture. A native MTP (nextn) head is welded on for single stream acceleration. ⚠️ Distilling a thinking style ≠ distilling knowledge/capability : the goal is "learn how to reason and how to converge ", not…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy