Dynin Omni: Omnimodal Unified Large Diffusion Language Model Introduction Unified masked diffusion modeling across textual reasoning, image generation, image editing, multi modal understanding, text to speech, and speech to text. Dynin Omni: Omnimodal Unified Large Diffusion Language Model is an 8B scale masked diffusion foundation model that unifies text, image, video, and speech understanding and generation within a single architecture. Unlike autoregressive (AR) unified models that serialize heterogeneous modalities into a left to right sequence, Dynin Omni models all modalities as discrete tokens in a shared vocabulary and performs generation via iterative masked denoising. This enables bidirectional context modeling, parallel multi token prediction, and globally conditioned any to any inference without modality specific expert decoders. Training proceeds in three stages: (1) modality adaptation, (2) omni modal supervised fine tuning with model merging, and (3) continual capability scaling. vLLM Omni Dynin Omni support has been merged into vLLM Omni through PR 1759 and is scheduled to be included in version 0.19.0 . Once 0.19.0 is released, this section will be updated with the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy