XPhoneBERT : A Pre trained Multilingual Model for Phoneme Representations for Text to Speech XPhoneBERT is the first pre trained multilingual model for phoneme representations for text to speech(TTS). XPhoneBERT has the same model architecture as BERT base, trained using the RoBERTa pre training approach on 330M phoneme level sentences from nearly 100 languages and locales. Experimental results show that employing XPhoneBERT as an input phoneme encoder significantly boosts the performance of a strong neural TTS model in terms of naturalness and prosody and also helps produce fairly high quality speech with limited training data. The general architecture and experimental results of XPhoneBERT can be found in our INTERSPEECH 2023 paper: @inproceedings{xphonebert, title = {{XPhoneBERT: A Pre trained Multilingual Model for Phoneme Representations for Text to Speech}}, author = {Linh The Nguyen and Thinh Pham and Dat Quoc Nguyen}, booktitle = {Proceedings of the 24th Annual Conference of the International Speech Communication Association (INTERSPEECH)}, year = {2023}, pages = {5506 5510} } Please CITE our paper when XPhoneBERT is used to help produce published results or is incorporated…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy