OmniVoice 🌍 OmniVoice is a massively multilingual zero shot text to speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model style architecture, it delivers high quality speech with superior inference speed, supporting voice cloning and voice design. Paper: OmniVoice: Towards Omnilingual Zero Shot Text to Speech with Diffusion Language Models Repository: GitHub Demo: Hugging Face Space Colab: Google Colab Notebook Key Features 600+ Languages Supported : The broadest language coverage among zero shot TTS models. Voice Cloning : State of the art voice cloning quality from a short reference audio. Voice Design : Control voices via assigned speaker attributes (gender, age, pitch, dialect/accent, whisper, etc.). Fine grained Control : Non verbal symbols (e.g., [laughter] ) and pronunciation correction via pinyin or phonemes. Fast Inference : RTF as low as 0.025 (40x faster than real time). Diffusion Language Model style Architecture : A clean, streamlined, and scalable design that delivers both quality and speed. Usage To get started, install the omnivoice library: We recommend using a fresh virtual environment (e.g.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy