Zonos v0.1 Zonos v0.1 is a leading open weight text to speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers. Our model enables highly natural speech generation from text prompts when given a speaker embedding or audio prefix, and can accurately perform speech cloning when given a reference clip spanning just a few seconds. The conditioning setup also allows for fine control over speaking rate, pitch variation, audio quality, and emotions such as happiness, fear, sadness, and anger. The model outputs speech natively at 44kHz. For more details and speech samples, check out our blog here We also have a hosted version available at playground.zyphra.com/audio Zonos follows a straightforward architecture: text normalization and phonemization via eSpeak, followed by DAC token prediction through a transformer or hybrid backbone. An overview of the architecture can be seen below. Usage Python Gradio interface (recommended) This should produce a sample.wav file in your project root directory. For repeated sampling we highly recommend using the gradio interface instead, as the minimal…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy