SeamlessM4T Medium SeamlessM4T is a collection of models designed to provide high quality translation, allowing people from different linguistic communities to communicate effortlessly through speech and text. This repository hosts 🤗 Hugging Face's implementation of SeamlessM4T. You can find the original weights, as well as a guide on how to run them in the original hub repositories (large and medium checkpoints). 🌟 SeamlessM4T v2, an improved version of this version with a novel architecture, has been released here. This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks. SeamlessM4T v2 is also supported by 🤗 Transformers, more on it in the model card of this new version or directly in 🤗 Transformers docs. SeamlessM4T Medium covers: 📥 101 languages for speech input ⌨️ 196 Languages for text input/output 🗣️ 35 languages for speech output. This is the "medium" variant of the unified model, which enables multiple tasks without relying on multiple separate models: Speech to speech translation (S2ST) Speech to text translation (S2TT) Text to speech translation (T2ST) Text to text translation (T2TT) Automatic speech recognition…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy