VibeVoice: A Frontier Open Source Text to Speech Model This repository contains a copy of model weights obtained from ModelScope(microsoft/VibeVoice Large). The license for this model is the MIT License , which permits redistribution . My understanding of the MIT License, which is consistent with the broader open source community's consensus, is that it grants the right to distribute copies of the software and its derivatives. Therefore, I am lawfully exercising the right to redistribute this model. If you are a rights holder and believe this understanding of the license is incorrect, please submit a DMCA complaint to Hugging Face at dmca@huggingface.co VibeVoice is a novel framework designed for generating expressive, long form, multi speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text to Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy