Dia is a 1.6B parameter text to speech model created by Nari Labs. It was pushed to the Hub using the PytorchModelHubMixin integration. Dia directly generates highly realistic dialogue from a transcript . You can condition the output on audio, enabling emotion and tone control. The model can also produce nonverbal communications like laughter, coughing, clearing throat, etc. To accelerate research, we are providing access to pretrained model checkpoints and inference code. The model weights are hosted on Hugging Face. The model only supports English generation at the moment. We also provide a demo page comparing our model to ElevenLabs Studio and Sesame CSM 1B. (Update) We have a ZeroGPU Space running! Try it now here. Thanks to the HF team for the support :) Join our discord server for community support and access to new features. Play with a larger version of Dia: generate fun conversations, remix content, and share with friends. 🔮 Join the waitlist for early access. ⚡️ Quickstart This will open a Gradio UI that you can work on. or if you do not have uv pre installed: Note that the model was not fine tuned on a specific voice. Hence, you will get different voices every time you…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy