Dataset Card for Unsupervised Peoples Speech Table of Contents Dataset Card for Unuspervised Peoples Speech Table of Contents Dataset Description Dataset Summary Dataset Structure Relevant Statistics Dataset Creation Source Data Annotations Considerations for Using the Data Additional Information Licensing Information Citation Information Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC BY and CC BY SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio files that have been stored as tar files, each containing a set of audio files. On average, each tar file is 5GB in size. All tar files are stored in either in the audio or audio2 directories. The licenses.jsonl file contains the license information for each audio file. The lang id results.jsonl file contains the predicted language for all files using Whisper Large V3. The vad results.jsonl file containes timestamps where voice was detected using Silero…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy