Qwopus3.6 27B v2 AWQ 4bit AutoAWQ format INT4 (W4A16) quantization of Jackrong/Qwopus3.6 27B v2, a Claude Opus reasoning distilled fine tune of Qwen 3.6 27B. The hybrid DeltaNet + softmax attention architecture is preserved, the 1 layer MTP head is included for speculative decoding, and the multimodal processor metadata is kept intact. APEX style edge protection keeps the first and last layers in BF16 for quality. Quick start Requires vLLM ≥ 0.21.0 : Benchmarks Evaluated with lm evaluation harness on a single NVIDIA B300 SXM6, 100 samples per task, 0 shot CoT, max gen toks=4096 : Task Qwen 3.6 27B (base) Qwopus 3.6 v2 (source BF16) This (AWQ 4bit) : : : GSM8K (flexible extract) 65.0% 87.0% 85.0% ARC Challenge (acc norm) 46.0% 45.0% 47.0% TruthfulQA MC2 55.1% 59.3% 59.3% IFEval (inst level strict) 40.5% 42.3% 42.9% Quantization preserves accuracy within standard error of the BF16 source on every task, and matches the source on TruthfulQA. The Claude Opus reasoning gain over the Qwen 3.6 base (+20 pp on GSM8K) is retained. Throughput Measured on a single NVIDIA B300 SXM6 with vLLM 0.21.0 and torch.compile enabled: Setup Throughput Speedup : : Batch = 1, no MTP 115 tok/s 1.00× Batch =…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy