Updates 2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k) README This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech. This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human For more details about YODAS dataset, please refer to our paper Usage: Considering the extremely large size of the entire dataset, we support two modes of dataset loadings: standard mode : each subset will be downloaded to the local dish before first iterating. streaming mode most of the files will be streamed instead of downloaded to your local deivce. It can be used to inspect this dataset quickly. Subsets/Shards There are 149 languages in this dataset, each language is sharded into at least 1 shard to make it easy for our processing and uploading purposes. The raw data of each shard contains 500G at most. Statistics of each shard can be found in the last section. We distinguish manual caption subset and automatic caption subset by the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy