Huihui Qwen3.6 27B abliterated NVFP4 MTP NVFP4 (modelopt W4A4 ) quant of Huihui Qwen3.6 27B abliterated — a Qwen3.5 family hybrid model (linear attention + periodic full attention) with a built in MTP (multi token prediction) head for speculative decoding. Multimodal capable ( Qwen3 5ForConditionalGeneration , vision/video tokens) but served here as a text generation / reasoning + tool calling model. Fits 4× 16 GB Blackwell (SM120) . ~7.2 GiB/GPU weights at TP=4 · 64K–262K context · reasoning · XML tool calls Single stream ~81–83 tok/s (TP=4, MTP n=3); peak ~880 tok/s @ 24 concurrent (64K) TL;DR — run it (no build required) The official vLLM image already ships the qwen3 5 architecture and the Qwen3 5MTP draft module, so you do not need to build anything. Or the raw docker run (what run.sh / compose.yaml wrap): Smoke test: Hardware & requirements 4× NVIDIA RTX PRO 2000 Blackwell (16 GB each, SM120), PCIe (no NVLink). Docker + NVIDIA Container Toolkit. The pre built vllm/vllm openai:v0.22.0 image carries vLLM ≥0.22 with NVFP4/modelopt + FlashInfer FP4 kernels and the qwen3 5 + MTP code. TP=4 sharding is clean: heads 24, KV heads 4, hidden 5120, intermediate 17408 — all ÷4. Bare meta…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy