Waxal Datasets The WAXAL dataset is a large scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large Scale Multilingual African Language Speech Corpus. Table of Contents Dataset Description ASR Dataset TTS Dataset How to Use Dataset Structure ASR Data Fields TTS Data Fields Data Splits Dataset Curation Considerations for Using the Data Additional Information Citation Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text to Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and language technology for these underserved languages, and to serve as a repository for digital preservation. The Waxal datasets are collections acquired through partnerships with Makerere University, The University of Ghana, Digital Umuganda, Media Trust, Loud and Clear, and AIMS Senegal. Acquisition was funded by Google and the Gates Foundation under an agreement to make the dataset openly accessible. The Senegalese languages (Wolof and Pular) were provided by AIMS Senegal. ASR Dataset The Waxal ASR dataset is a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy