Quick Links Resource Link Model Weights + Full Documentation AEON 7/Gemma 4 26B A4B it Uncensored NVFP4 on HuggingFace DFlash vLLM Container (DGX Spark) ghcr.io/aeon 7/aeon vllm ultimate:latest DFlash Drafter z lab/gemma 4 26B A4B it DFlash Quick Start Recipe notes — the drafter backend is tied to the image tag: On :latest / :2026 06 11 pr41703 (current): the drafter must use attention backend: flash attn (enabled for multimodal Gemma targets by PR 41703's use mm prefix=False ). flex attention crashes at the first request on this image (non contiguous KV view in the new KV sharing path). enable prefix caching is safe and recommended — soak validated with DFlash (see the fixes section below). On the rollback tag :2026 06 04 pr44389 (pre fix): the reverse — only flex attention loads ( flash attn fails with partial multimodal token full attention not supported ), the drafter silently runs without sliding window support (~0% acceptance beyond 2k token contexts), and enable prefix caching triggers a slow acceptance collapse (see below). Not recommended. Either way the body stays on triton attn (Gemma's heterogeneous head dims), and do not set kv cache dtype fp8 — DFlash's non causal dra…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy