Qwen3.6 27B AEON Ultimate Uncensored FP8 MTP A vLLM ready FP8 quantization of AEON 7/Qwen3.6 27B AEON Ultimate Uncensored (an abliterated fine tune of Qwen3.6 27B), packaged in block 128 FP8 with the Multi Token Prediction (MTP) draft head taken verbatim from Qwen/Qwen3.6 27B FP8 for vLLM speculative decoding. In numbers: 0/100 refusals on mlabonne/harmful behaviors[:100] (vs 100/100 for vanilla Qwen3.6 27B FP8) — abliteration preserved +1–3 pp on gsm8k / ifeval vs vanilla — capability preserved +90 % decode TPS vs the same checkpoint with no MTP, via speculative decoding at K=3 (~43–45 TPS on RTX A6000 / Ampere) Byte shape compatible with Qwen/Qwen3.6 27B FP8 — quant method: "fp8" , weight block size: [128, 128] , single vLLM Fp8LinearMethod loader path What's in the box Component Format Body weights (Linear modules outside the exclusion list) FP8 e4m3fn, block 128 ( weight scale inv , shape (out/128, in/128) ) Vision tower, lm head , embed tokens , linear attn.in proj {a,b,ba} SSM state projections BF16 (matches Qwen's modules to not convert list, 882 entries) MTP block ( mtp. ) Verbatim from Qwen/Qwen3.6 27B FP8 — 7 FP8 attention/MLP weights with block 128 scales + 8 BF16 norms…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy