MOVA: Towards Scalable and Synchronized Video–Audio Generation We introduce MOVA ( MO SS V ideo and A udio), a foundation model designed to break the "silent era" of open source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment. 🌟Key Highlights Native Bimodal Generation : Moves beyond clunky cascaded pipelines. MOVA generates high fidelity video and synchronized audio in a single inference pass, eliminating error accumulation. Precise Lip Sync & Sound FX : Achieves state of the art performance in multilingual lip synchronization and environment aware sound effects. Fully Open Source : In a field dominated by closed source models (Sora 2, Veo 3, Kling), we are releasing model weights, inference code, training pipelines, and LoRA fine tuning scripts. Asymmetric Dual Tower Architecture : Leverages the power of pre trained video and audio towers, fused via a bidirectional cross attention mechanism for rich modality interaction. Demo Model Details Model Description MOVA addresses the limitations of proprietary systems like Sora 2 and Veo 3 by offering a fully open source framework fo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy