AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large scale, multilingual dataset designed for a wide range of audio and speech processing tasks. It comprises a diverse collection of audio clips, including both spoken voice and music, making it a valuable resource for training and evaluating models for automatic speech recognition (ASR), text to speech (TTS), audio classification, and more. The voice data is aggregated from well known public corpora such as Emilia , LibriTTS R , and Common Voice . The music portion is sourced from various publicly available datasets. To ensure comprehensive and consistent annotation, the dataset has been enhanced with state of the art AI models: Transcriptions : Missing transcriptions for voice entries were generated using OpenAI's Whisper model. Descriptions : Descriptive metadata for audio content was generated using the Qwen2 Audio model. Dataset Structure Data Instances A typical example from the dataset looks like this: Data Fields The dataset contains the following fields: id (string): A unique identifier for each audio clip. description (string): A textual description of the audio content. Generated by Qwen2.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy