Thanks to @aoi ot for HF reupload, this is duplicated from their repo :) https://github.com/vibevoice community/VibeVoice VibeVoice: A Frontier Open Source Text to Speech Model VibeVoice is a novel framework designed for generating expressive, long form, multi speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text to Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next token diffusion framework, leveraging a Large Language Model (LLM) to understand textual context and dialogue flow, and a diffusion head to generate high fidelity acoustic details. The model can synthesize speech up to 90 minutes long with up to 4 distinct speakers , surpassing the typical 1 2 speaker limits of many prior models. ➡️ Technical Report: VibeVoice Technical Report ➡️ Project Page: microsoft/VibeVoic…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy