qwen3.5 4B super coder (Q 4.0 GGUF) qwen3.5 4B super coder is a 4 bit quantized GGUF model optimized for fast, reliable coding, structured tool calling, and active reasoning (thinking mode) on consumer/mobile hardware. It is distilled from Claude Sonnet 4.6 & Opus 4.6, and merged/quantized using Unsloth. Model Summary & Architecture Base Model : Qwen/Qwen3.5 4B Format : GGUF (Q 4.0 Quantization) Size : ~2.6 GB Context Window : 32K (optimized for mobile RAM budgets, natively supports up to 262K/1M context via YaRN) Key Architectural Advantage : The base Qwen3.5 4B model uses a hybrid architecture combining Gated DeltaNet (3 layers) and Full Attention (1 layer) repeating. Since only 8 of the 32 layers store a full KV cache, the KV cache footprint is incredibly small (~0.4GB for 32K context), making it exceptionally well suited for high context coding on mobile devices (e.g., iPhone 15 Pro+, flagship Android, iPad Pro). Distillation & Training Procedure This model was trained using a staged Supervised Fine Tuning (SFT) pipeline to systematically inject reasoning capability, coding specialization, and tool calling precision: 1. Phase 1: Distillation (Claude Behavior) Dataset : clzoro/C…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy