🎵 AudioMarathon: A Comprehensive Benchmark for Long Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large scale, multi task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long form audio content. It provides a diverse set of 10 tasks built upon three pillars: long context audio inputs with durations ranging from 90.0 to 300.0 seconds, which correspond to encoded sequences of 2,250 to 7,500 audio tokens, respectively, full domain coverage across speech, sound, and music, and complex reasoning that requires multi hop inference. 📊 Task Taxonomy & Statistics Task Categories AudioMarathon organizes tasks into four primary categories: 1. Speech Content Extraction 2. Audio Classification 3. Speaker Information Modeling Dataset Statistics Task ID Dataset Task Type Samples Duration Format License Status 1 LibriSpeech long Automatic Speech Recognition (ASR) 204 1 4min FLAC 16kHz CC BY 4.0 ✅ Full 2 RACE Speech Content Reasoning (SCR) 820 2 4.22min WAV 16kHz Apache 2.0 ✅ Full 3 HAD Speech Detection (SD) 776 3~5min WAV 16kHz CC BY 4.0 ✅ Full 4 GTZAN Music c…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy