license: cc by nc 4.0 tags: mms vits pipeline tag: text to speech Massively Multilingual Speech (MMS): Spanish Text to Speech This repository contains the Spanish (spa) language text to speech (TTS) model checkpoint. This model is part of Facebook's Massively Multilingual Speech project, aiming to provide speech technology across a diverse range of languages. You can find more details about the supported languages and their ISO 639 3 codes in the MMS Language Coverage Overview, and see all MMS TTS checkpoints on the Hugging Face Hub: facebook/mms tts. MMS TTS is available in the 🤗 Transformers library from version 4.33 onwards. Model Details VITS ( V ariational I nference with adversarial learning for end to end T ext to S peech) is an end to end speech synthesis model that predicts a speech waveform conditional on an input text sequence. It is a conditional variational autoencoder (VAE) comprised of a posterior encoder, decoder, and conditional prior. A set of spectrogram based acoustic features are predicted by the flow based module, which is formed of a Transformer based text encoder and multiple coupling layers. The spectrogram is decoded using a stack of transposed convolutio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy