Dataset Card for Voxpopuli Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://github.com/facebookresearch/voxpopuli Repository: https://github.com/facebookresearch/voxpopuli Paper: https://arxiv.org/abs/2101.00390 Point of Contact: changhan@fb.com, mriviere@fb.com, annl@fb.com Dataset Summary VoxPopuli is a large scale multilingual speech corpus for representation learning, semi supervised learning and interpretation. The raw data is collected from 2009 2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non native English intended for rese…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy