Qwen3.6 35B A3B heretic — NVFP4 (v2 multimodal preserved) ✅ Validated 2026 06 09 on the unified AEON vLLM Ultimate image ghcr.io/aeon 7/aeon vllm ultimate:latest (vLLM 0.22.1+pr44389) — loads + serves cleanly (compressed tensors MoE) — measured 40.6 tok/s single / 289 tok/s conc×16 (no spec decode). Recommended container base. 🚀 PRODUCTION DEPLOYMENT GUIDE: github.com/AEON 7/Qwen3.6 NVFP4 DFlash The GitHub repo is the definitive turn key setup for DGX Spark — pre built Docker image, end to end deployment guide, validated OpenClaw config, the 8 vLLM patches that actually make this work on SM121, and a concurrency sweep benchmark harness. Image: ghcr.io/aeon 7/vllm spark omni q36:v1.2 (vLLM HEAD source built for cu130/sm 120 + 5 source patches + flashinfer 0.6.8 + Marlin GEMM enforcement) Pairs with z lab/Qwen3.6 35B A3B DFlash drafter (must be post 2026 04 19 revision) Production stable under sustained chat load — measured 116.8 tok/s single stream / 785.3 tok/s aggregate at 128 concurrency on DGX Spark What changed in v2 (2026 04 19) v1 of this checkpoint had model.language model.layers.X. keys remapped to model.layers.X. so vLLM's text only Qwen3 5MoeForCausalLM loader would pick…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy