ⓍTTS ⓍTTS is a Voice generation model that lets you clone voices into different languages by using just a quick 6 second audio clip. There is no need for an excessive amount of training data that spans countless hours. This is the same or similar model to what powers Coqui Studio and Coqui API. Features Supports 17 languages. Voice cloning with just a 6 second audio clip. Emotion and style transfer by cloning. Cross language voice cloning. Multi lingual speech generation. 24khz sampling rate. Updates over XTTS v1 2 new languages; Hungarian and Korean Architectural improvements for speaker conditioning. Enables the use of multiple speaker references and interpolation between speakers. Stability improvements. Better prosody and audio quality across the board. Languages XTTS v2 supports 17 languages: English (en), Spanish (es), French (fr), German (de), Italian (it), Portuguese (pt), Polish (pl), Turkish (tr), Russian (ru), Dutch (nl), Czech (cs), Arabic (ar), Chinese (zh cn), Japanese (ja), Hungarian (hu), Korean (ko) Hindi (hi) . Stay tuned as we continue to add support for more languages. If you have any language requests, feel free to reach out! Code The code base supports inferen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy