LAION Audio 630K Freesound Dataset LAION Audio 630K is the largest audio text dataset publicly available and a magnitude larger than previous audio text datasets (by 2022 11 05). Notably, it combines eight distinct datasets, which includes the Freesound dataset. Specifically, this Hugging face repository contains two versions of Freesound dataset. Details of each dataset (e.g. how captions are made etc.) could be found in the "datacard" column of the table below. Freesound (full) : The complete Freesound dataset, available at /freesound folder. Freesound (no overlap) : Made based on Freesound(full), with samples from ESC50, FSD50K, Urbansound8K and Clotho removed. available at /freesound no overlap folder. As of the structure and format of freesound and freesound no overlap folder, please refer to this page. Name Duration Number of Samples Data Type Metadata Data Card Freesound (no overlap) 2817.31hrs 460801 1 2 captions per audio, audio website csv data card Freesound (full) 3033.38hrs 515581 1 2 captions per audio, audio website csv data card Metadata csv file For each of the two datasets, we provide a metadata csv file including the following columns: audio filename : The filena…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy