Model Overview Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Description: Audio Flamingo 3 (AF3) is a fully open, state of the art Large Audio Language Model (LALM) that advances reasoning and understanding across speech, sounds, and music. AF3 builds on previous work with innovations in: Unified audio representation learning (speech, sound, music) Flexible, on demand chain of thought reasoning Long context audio comprehension (up to 10 minutes) Multi turn, multi audio conversational dialogue (AF3 Chat) Voice to voice interaction (AF3 Chat) Extensive evaluations confirm AF3’s effectiveness, setting new benchmarks on over 20 public audio understanding and reasoning tasks. This model is for non commercial research purposes only. Usage Audio Flamingo 3 is supported in 🤗 Transformers. To run the model, first install Transformers: Note: AF3 processes audio in 30 second windows with a 10 minute total cap per sample. Longer inputs are truncated. Single turn: audio + text instruction Multi turn chat Batch multiple conversations Text only and audio only prompts AF3 transcription checkpoints prepend answers with fixed assistant phrasing such as T…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy