Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled — AWQ INT4 (W4A16, multimodal) INT4 quantization (AutoAWQ GEMM layout, W4A16, group size 128, asymmetric) of lordx64/Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled , a 35B parameter / 3B active Qwen3.5 MoE multimodal model that has been reasoning distilled from Claude 4.7 Opus. This artifact preserves the full multimodal stack (vision tower, linear attention layers, shared experts, layer 0, MTP head, router gates) at fp16/bf16 and only quantizes the routed expert MLPs ( gate proj / up proj / down proj per expert × 256 experts × ~39 routed MoE layers). The on disk layout mirrors QuantTrio/Qwen3.6 35B A3B AWQ exactly so vLLM auto detects and runs it with the same moe wna16 / AWQ kernel path. If you don't need multimodal and want maximum quality, use the sibling FP8 build — feanorscode/Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled FP8 Dynamic . Method note — read this before citing as "AWQ". This build does data free RTN quantization packed in the AutoAWQ GEMM format . It uses the AWQ on disk layout and the AWQ aware vLLM kernel path, but does not run AWQ's defining activation aware salience pass over a calibration corpus.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy