Model Overview Music Flamingo: Scaling Music Understaning in Audio Language Models Description: Music Flamingo (MF) is a fully open, state of the art Large Audio Language Model (LALM) designed to advance music (including song) understanding in foundational audio models. MF brings together innovations in: Deep music understanding across songs and instrumentals. Rich, theory aware captions and question answering (harmony, structure, timbre, lyrics, cultural context). Reasoning centric training using chain of thought + reinforcement learning with custom rewards for step by step reasoning. Long form song reasoning over full length, multicultural audio (extended context). Extensive evaluations confirm Music Flamingo's effectiveness, setting new benchmarks on over 10+ public music understanding and reasoning tasks. This model is for non commercial research purposes only. Usage Music Flamingo (MF) is supported in 🤗 Transformers. To run the model, first install Transformers from our fork (we are awaiting HF merge): Note: MF processes audio in 30 second windows with a 20 minute total cap per sample. Longer inputs are truncated. ➡️ audio + text instruction ➡️ multi turn: ➡️ text only: ➡️ au…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy