Qwen3.5 27B W4A16 AWQ HD Agent This is a highly optimized, hybrid W4A16 AWQ quantization of Qwen/Qwen3.5 27B. It was built using a custom llmcompressor pipeline designed to preserve maximum reasoning, agentic logic, and structural integrity. 🎯 Hybrid AWQ Quantization Strategy Unlike standard uniform 4 bit AWQ this model uses a surgical, hybrid precision recipe. Crucial structural and recurrent layers were explicitly bypassed to prevent the logic degradation and multimodal failure typical of heavy quantization: Protected Layers (16 bit): lm head , embed tokens , linear attn (recurrent state preservation), vision towers ( model.visual. ), and next token prediction blocks ( mtp. ). Targeted Smoothing: AWQ scaling was selectively applied using a custom mapped layer balance (smoothing the input layernorm of every 4th attention layer, and all MLP blocks universally) to maintain coherence across extended generations. Weight Precision: 4 bit INT Group Size: 128 📚 High Depth Calibration Mixture To anchor the model's performance for agentic workflows, the calibration process utilized a precisely balanced 1,024 sample dataset at a 4096 sequence length. The mixture is heavily skewed toward r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy