MedDialogue Audio English Medical Dialogue Corpus for Speech Recognition Research. This repository contains MedDialogue Audio, an English audio corpus designed for research in Automatic Speech Recognition (ASR) in the healthcare domain. The dataset was published in the proceedings of the 7th SBBD Dataset Showcase Workshop, and is available online at the following link: https://sol.sbc.org.br/index.php/dsw/article/view/37199 Dataset Description MedDialogue Audio is derived from the MedDialog EN transcription dataset. It aims to support the development and evaluation of ASR systems under acoustic conditions that simulate clinical environments. The creation process consisted of three main steps: 1. Text Normalization The transcriptions from the original corpus were processed using a language model to perform corrections and standardization. 2. Speech Synthesis The normalized texts were converted into audio using a Text to Speech (TTS) model. 3. Acoustic Data Augmentation Variants of the audio files were generated by adding white noise and hospital background sounds at multiple intensity levels. The final corpus consists of 10,534 dialogues, resulting in a total of 147,476 audio files.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy