Speech Recognition Alignment Dataset This dataset is a variation of several widely used ASR datasets, encompassing Librispeech, MuST C, TED LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: The dataset can be downloaded and pre processed on disk using the load dataset function: It can also be streamed directly from the Hub using Datasets' streaming mode. Loading a dataset in streaming mode loads individual samples of the dataset at a time, rather than downloading the entire dataset to disk: Citation If you use this data, please consider citing the ICASSP 2024 Paper: SYNTHETIC CONVERSATIONS IMPROVE MULTI TALKER ASR: License This dataset is licensed in accordance with the terms of the original dataset.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy