Qwen3.6 27B AEON Ultimate Uncensored Multimodal NVFP4 MTP XS Deployment, operations & benchmarks → github.com/AEON 7/Qwen3.6 27B AEON Ultimate Uncensored DFlash The GitHub repo is the source of truth for the production deployment guide, hardware tuned docker compose configs, full configuration reference, measured benchmarks, and AGENTS.md — an operator's manual that pre empts common stale documentation traps. DGX Spark performance — previous production (v3 image, 2026 04 29 — superseded by the DFlash n=12 recipe below) Served with DFlash spec decode (not the MTP head) on this XS body, the v3 image ( ghcr.io/aeon 7/vllm aeon ultimate dflash:qwen36 v3 ) clocks 38.5 tok/s median, 71.3 tok/s peak thinking on / 38.1 / 68.4 thinking off — a +18 % median / +26 % peak lift over the prior v2.1 image and a +17 % / +21 % stacked lift vs the original NVFP4 (compressed tensors) production. Median TTFT is 247 ms (was 325 ms — −24 %). See the GitHub Performance section for the four config comparison table. 🆕 Next gen container — AEON vLLM Ultimate (2026 06 04) ghcr.io/aeon 7/aeon vllm ultimate:latest — vLLM 0.22.1 + PR 44389 Triton NVFP4 KV cache (~3× KV capacity) + DFlash + TurboQuant K8V4 + AE…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy