[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ StrongAC GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain of Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly. Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step by step temporal analysis. 🏆 DCASE 2026 Challenge: We are proud to announce that AudioMCQ StrongAC GeminiCoT serves as the official training set for DCASE 2026 Challenge Task 5, Audio Dependent Question Answering. News 2025 03 31 : Added more data and filtered out low quality CoT samples (e.g., hallucinations claiming no audio access or visual hallucinations). Now contains 19,480 samples. 2025 03 29 : Initial release with 16,407 samples. Dataset Origin This dataset builds upon the foundation of our ICLR 2026 paper, "Measuring Audio's Impact on Correctness: Audio Contribution Aware Post Training of Large Audio Language Models" . 1. Base Dataset : AudioMCQ's Strong Audio Contribution (StrongAC) split. Construction (Multi model Consensus) : Tested 3 LALMs usi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy