Qwen3.6 27B AEON Ultimate Uncensored Text NVFP4 MTP XS Deployment, operations & benchmarks → github.com/AEON 7/Qwen3.6 27B AEON Ultimate Uncensored DFlash The GitHub repo is the source of truth for the production deployment guide, hardware tuned docker compose configs, full configuration reference, measured benchmarks, and AGENTS.md — an operator's manual that pre empts common stale documentation traps. 🏆 DGX Spark performance — current production (v3 image, 2026 04 29) The XS body served with DFlash spec decode (not the MTP head) under the v3 image ( ghcr.io/aeon 7/vllm aeon ultimate dflash:qwen36 v3 ) is the highest throughput config we've measured on Spark: 38.5 tok/s median, 71.3 tok/s peak thinking on / 38.1 / 68.4 thinking off. That's a +17–26 % lift across thinking modes vs the original NVFP4 + old DFlash + v2.1 production. See the GitHub Performance section for the four config comparison table. 🙏 Reference recipe credit: The conv1d preserved NVFP4 + MTP graft pipeline used to build this XS variant is based on sakamakismile 's validated Qwen3.6 27B NVFP4 MTP series (22K+ downloads). They worked out the modelopt config — including the strategic decision to quantize the GDN…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy