WavCaps WavCaps is a ChatGPT assisted weakly labelled audio captioning dataset for audio language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly labelled Subset). Paper: https://arxiv.org/abs/2303.17395 Github: https://github.com/XinhaoMei/WavCaps Statistics Data Source audio avg. audio duration (s) avg. text length FreeSound 262300 85.98 6.77 BBC Sound Effects 31201 115.04 9.67 SoundBible 1232 13.12 5.87 AudioSet SL subset 108317 10.00 9.79 WavCaps 403050 67.59 7.80 Download We provide a json file for each data source. For audio clips sourced from websites, we provide processed caption, raw description, as well as other metadata. For audio clips from AudioSet, we use the version from PANNs, where each file name is appended with a 'Y' at the start. For the start time, please refer to the original metadata of AudioSet SL subset. Waveforms with flac format can be downloaded through Zip files directory. Pretrained models can be downloaded here. If you get "error: invalid zip file with overlapped components (possible zip bomb)" when unzipping, please try the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy