Voices in the Wild Project Page Paper GitHub Voices in the Wild (Voices in the Wild 2M) is a large scale automatic speech recognition (ASR) dataset designed for robustness training and evaluation under diverse, real world acoustic conditions. It covers 7 classic acoustic phenomena (including noise, far field speech, obstruction, echo/reverberation, recording artifacts, electronic distortion, and transmission dropout) and 54 physically plausible compound scenarios. The dataset was introduced as part of the Mega ASR framework to address the "acoustic robustness bottleneck" where models produce omissions or hallucinations under severe compositional distortions. Data Fields file name : relative path to the audio file. audio path : audio path retained for local tooling. text : transcription alias copied from answer . answer : reference transcription. question : transcription instruction. subset : normalized acoustic condition category. prediction : empty placeholder for model output. name : public sample identifier. index : integer sample index. Dataset Size Total examples : 645,925 Subset categories : 54 Loading Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy