gemma 4 31B it DFlash Paper GitHub Blog DFlash is a speculative decoding method that uses a lightweight block diffusion model to draft multiple tokens in parallel. This is the drafter model, which must be paired with google/gemma 4 31B it. Quick Start Installation vLLM: until Gemma4 DFlash support is merged, install vLLM from PR 41703: SGLang: Launch Server vLLM: SGLang: Usage For vLLM, use port 8000 . For SGLang, use port 30000 . Benchmark Results Setup: Single NVIDIA B300 GPU per server/run, vLLM, thinking enabled, max output length 4096, greedy decoding. Throughput and Speedup DFlash achieves up to 5.8x speedup at concurrency 1. Generated tokens/sec (speedup vs. autoregressive baseline) Block Size = 16 Task Concurrency AR DFlash : : : Math500 1 77 447 (5.8x) 8 511 2650 (5.2x) 32 1308 4962 (3.8x) GSM8K 1 78 408 (5.3x) 8 520 2321 (4.5x) 32 1382 4447 (3.2x) HumanEval 1 76 420 (5.6x) 8 494 2389 (4.8x) 32 1145 4139 (3.6x) MBPP 1 79 343 (4.4x) 8 535 2036 (3.8x) 32 1389 3636 (2.6x) MT Bench 1 79 236 (3.0x) 8 503 1334 (2.7x) 32 1177 2257 (1.9x) Acceptance Length Task c1 c8 c32 : : : Math500 8.59 8.59 8.62 GSM8K 7.53 7.50 7.52 HumanEval 8.00 7.89 7.96 MBPP 6.13 6.13 6.14 MT Bench 4.23 4.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy