Qwen3.6 27B AEON Ultimate Uncensored BF16 Deployment, operations & benchmarks → github.com/AEON 7/Qwen3.6 27B AEON Ultimate Uncensored DFlash The GitHub repo is the source of truth for the production deployment guide, hardware tuned docker compose configs (DGX Spark NVFP4, A100/H100 BF16), full configuration reference, measured throughput benchmarks, and AGENTS.md — an operator's manual that pre empts common stale documentation traps for AI coding agents working on this stack. 🆕 2026 05 01 — MTP head grafted in. This repo now ships with the original mtp. head (15 tensors, ~0.85 GB) restored from the Qwen/Qwen3.6 27B base. vLLM's speculative config '{"method":"qwen3 5 mtp","num speculative tokens":3}' works directly on the BF16 checkpoint with no extra steps. Measured on DGX Spark: mean accepted length 3.3/3, P0 ≈ 90% acceptance, avg draft acceptance 78% — comparable to the base model, confirming abliteration doesn't damage the model's top K distribution that MTP relies on. Credit to @tcclaviger for the empirical finding (discussion 6) that MTP doesn't need retraining post abliteration. No retraining was performed. The MTP head is the unmodified base — only the residual stream writ…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy