━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo Audio: Audio Language Models are Few Shot Learners ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 🤖 GitHub 📄 Paper 📰 Blog 🔥 Online Demo 📊 MiMo Audio Eval Introduction Existing audio language models typically rely on task specific fine tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT 3 has shown that scaling next token prediction pretraining enables strong generalization capabilities in text, and we believe this paradigm is equally applicable to the audio domain. By scaling MiMo Audio's pretraining data to over one hundred million of hours, we observe the emergence of few shot learning capabilities across a diverse set of audio tasks. We develop a systematic evaluation of these capabilities and find that MiMo Audio 7B Base achieves SOTA performance on both speech intelligence and audio understanding benchmarks among open source models. Beyond standard metrics, MiMo Audio 7B Base generalizes to tasks absent from its training data, such as voice conversion, style transfer, and speech editing. Mi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy