Qwopus3.6 27B v2 NVFP4 Mixed precision (NVFP4 + FP8 + BF16) quantization of Jackrong/Qwopus3.6 27B v2, a Claude Opus reasoning distilled fine tune of Qwen 3.6 27B. The hybrid DeltaNet + softmax attention architecture is preserved, the 1 layer MTP head is included as a BF16 sidecar for speculative decoding, and the multimodal processor metadata is kept intact. Quick start Requires vLLM ≥ 0.21.0 and a Blackwell class GPU (SM 10.0+) for native NVFP4 W4A4 inference: Benchmarks Evaluated with lm evaluation harness on a single NVIDIA B300 SXM6, 100 samples per task, 0 shot CoT, max gen toks=4096 : Task Qwen 3.6 27B (base) Qwopus 3.6 v2 (source BF16) This (NVFP4) : : : GSM8K (flexible extract) 65.0% 87.0% 87.0% ARC Challenge (acc) 50.0% 50.0% 53.0% TruthfulQA MC2 55.1% 59.3% 58.7% IFEval (inst level strict) 40.5% 42.3% 41.7% Accuracy is preserved versus the BF16 source — the GSM8K score is identical to the source and the other tasks match within standard error. Throughput Measured on a single NVIDIA B300 SXM6 with vLLM 0.21.0 and torch.compile enabled: Setup Throughput Speedup : : Batch = 1, no MTP 121 tok/s 1.00× Batch = 1, MTP num speculative tokens = 3 274 tok/s 2.26× Batch = 8 continu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy