Marco LongSpeech Dataset Marco LongSpeech is a multi task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs. 📊 Dataset Statistics Task Statistics Task Train Val Test Total Unique Audios ASR 71,275 15,273 15,274 101,822 101,822 Temporal Relative QA 5,886 1,261 1,262 8,409 8,409 summary 4,366 935 937 6,238 6,238 content separation 5,887 1,261 1,263 8,411 8,411 emotionQA 5,887 1,261 1,263 8,411 8,411 speaker count 5,887 1,261 1,263 8,411 8,411 translation 29,435 6,307 6,309 42,051 8,411 language detection 14,789 3,169 3,170 21,128 21,128 Total 143,412 30,728 30,741 204,881 Audio Subset Statistics Subset WAV Files all audios.jsonl metadata.json LongSpeech p1 29,539 ✓ ✓ LongSpeech p2 22,107 ✓ ✓ LongSpeech p3 50,176 ✓ ✓ Total 101,822 📁 Dataset Structure 🎯 Task Descriptions The dataset covers a comprehensive range of capabilities required for long speech understanding: ASR & S2T Translation : Core transcription and translation of full length audio. Summarization : Generating concise summaries from lengthy recordings. Speaker Count & Language Detection : Identifying speaker and language…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy