Nemotron Labs Diffusion 8B Model Overview Nemotron Labs Diffusion is a tri mode language model that supports both AR decoding and diffusion based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self speculation: the same model performs diffusion based parallel drafting and AR verification with shared KV cache, achieving high acceptance lengths and decoding efficiency. The seamless mode switching by simply changing attention patterns enables high efficiency at different concurrency levels in varying deployment scenarios with one single model. Highlights SOTA 3B, 8B, 14B dense LM family (base, instruct, and vision language variants) supporting AR, diffusion, and self speculation with the focus on decode efficiency. Generation moved from a memory bound regime toward a compute bound regime. Model weights are loaded once and reused to compute multiple tokens during generation. Self speculation uses diffusion for drafting and AR for verification, providing a stronger alternative to MTP approaches: 3x higher acceptance length and 2.2x speed up vs. Qwen3 8B Eagle3 in SGLang. 5.9…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy