InfiniteTalk: Audio driven Video Generation for Sparse Frame Video Dubbing We propose InfiniteTalk , a novel sparse frame video dubbing framework. Given an input video and audio track, InfiniteTalk synthesizes a new video with accurate lip synchronization while simultaneously aligning head movements, body posture, and facial expressions with the audio. Unlike traditional dubbing methods that focus solely on lips, InfiniteTalk enables infinite length video generation with accurate lip synchronization and consistent identity preservation. Beside, InfiniteTalk can also be used as an image audio to video model with an image and an audio as input. 💬 Sparse frame Video Dubbing – Synchronizes not only lips, but aslo head, body, and expressions ⏱️ Infinite Length Generation – Supports unlimited video duration ✨ Stability – Reduces hand/body distortions compared to MultiTalk 🚀 Lip Accuracy – Achieves superior lip synchronization to MultiTalk This repository hosts the model weights for InfiniteTalk . For installation, usage instructions, and further documentation, please visit our GitHub repository. License Agreement The models in this repository are licensed…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy