Higgs TTS 2: Redefining Expressiveness in Audio Generation Check our open source repository https://github.com/boson ai/higgs audio for more details! Rename note: Higgs Audio V2 and Higgs Audio V2 Generation have been renamed to Higgs TTS 2. We are open sourcing Higgs TTS 2, a powerful audio foundation model pretrained on over 10 million hours of audio data and a diverse set of text data. Despite having no post training or fine tuning, Higgs TTS 2 excels in expressive audio generation, thanks to its deep language and acoustic understanding. On EmergentTTS Eval, the model achieves win rates of 75.7% and 55.7% over "gpt 4o mini tts" on the "Emotions" and "Questions" categories, respectively. It also obtains state of the art performance on traditional TTS benchmarks like Seed TTS Eval and Emotional Speech Dataset (ESD). Moreover, the model demonstrates capabilities rarely seen in previous systems, including automatic prosody adaptation during narration, zero shot generation of natural multi speaker dialogues in multiple languages, melodic humming with the cloned voice, and simultaneous generation of speech and background music. Here's the demo video that shows some of its emergent cap…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy