π΅ Matxa TTS (Matcha TTS) Catalan Multiaccent Table of Contents Click to expand Model description Intended uses and limitations How to use Training Evaluation Citation Additional information Summary Here we present π΅ Matxa, the first multispeaker, multidialectal neural TTS model. It works together with the vocoder model π₯ alVoCat, to generate high quality and expressive speech efficiently in four dialects: Balear Central North Occidental Valencian Both models are trained with open data; π΅ Matxa models are free (as in freedom) to use for non comercial purposes, but for commercial purposes it needs licensing from the voice artist. To listen to the voices you can visit the dedicated space. Model Description π΅ Matxa TTS is based on Matcha TTS that is an encoder decoder architecture designed for fast acoustic modelling in TTS. The encoder part is based on a text encoder and a phoneme duration prediction that together predict averaged acoustic features. And the decoder has essentially a U Net backbone inspired by Grad TTS, which is based on the Transformer architecture. In the latter, by replacing 2D CNNs by 1D CNNs, a large reduction in memory consumption and fast synthesis is achieβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy