Base model: Qwen/Qwen3.5 4B Qwen3.5 4B , self quantized to GGUF by Atomic Chat. Built straight from Qwen's original weights with a per tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. Highlights 4.7B parameters : the weights this repo quantizes. Context length : 262,144 tokens (256K), as published by Qwen. 32 layers : Dense decoder. Modalities : the base model handles Text, Image; this repo ships text only quants, it carries no vision projector. Full imatrix ladder : every quant is calibrated with an importance matrix. Unified Vision Language Foundation : Early fusion training on multimodal tokens achieves cross generational parity with Qwen3 and outperforms Qwen3 VL models across reasoning, coding, agents, and visual understanding benchmarks. Efficient Hybrid Architecture : Gated Delta Networks combined with sparse Mixture of Experts deliver high throughput inference with minimal latency and cost overhead. [!NOTE] These GGUFs are self quantized from the original weights , not a repack. The importance matrix keeps low bit quants closer to the full precision model. [!IMPORTANT] Always pass jinja so the Qwen3.5 4B chat template is applie…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy