SeamlessM4T v2 SeamlessM4T is our foundational all in one M assively M ultilingual and M ultimodal M achine T ranslation model delivering high quality translation for speech and text in nearly 100 languages. SeamlessM4T models support the tasks of: Speech to speech translation (S2ST) Speech to text translation (S2TT) Text to speech translation (T2ST) Text to text translation (T2TT) Automatic speech recognition (ASR). SeamlessM4T models support: 🎤 101 languages for speech input. 💬 96 Languages for text input/output. 🔊 35 languages for speech output. 🌟 We are releasing SeamlessM4T v2, an updated version with our novel UnitY2 architecture. This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks. The v2 version of SeamlessM4T is a multitask adaptation of our novel UnitY2 architecture. Unity2 with its hierarchical character to unit upsampling and non autoregressive text to unit decoding considerably improves over SeamlessM4T v1 in quality and inference speed. SeamlessM4T v2 is also supported by 🤗 Transformers, more on it in the dedicated section below. SeamlessM4T models Model Name params checkpoint metrics SeamlessM4T Large v2 2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy