Dataset Card for People's Speech Table of Contents Dataset Description Dataset Summary Supported Tasks Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Dataset Description Homepage: https://mlcommons.org/en/peoples speech/ Repository: https://github.com/mlcommons/peoples speech Paper: https://arxiv.org/abs/2111.09344 Leaderboard: [Needs More Information] Point of Contact: datasets@mlcommons.org Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC BY SA and CC BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech to text systems and crucially is available with a permissive license. Supported Tasks and Leaderboards [Needs More Information] Languages English Dataset Str…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy