SpeechEval SpeechEval is a large scale multilingual dataset for general purpose, interpretable speech quality evaluation , introduced in the paper: SpeechLLM as Judges: Towards General and Interpretable Speech Quality Evaluation It is designed to train and evaluate Speech LLMs acting as “judges” that can explain their decisions, compare samples, suggest improvements, and detect deepfakes. 1. Dataset Overview Utterances: 32,207 unique speech clips Annotations: 128,754 human verified annotations Languages: English, Chinese, Japanese, French Modalities: Audio + Natural language annotations License: CC BY NC SA 4.0 Each example combines structured labels and rich natural language explanations , making it suitable for both classic supervised learning and instruction tuning of SpeechLLMs. The dataset covers four core evaluation tasks : 1. Speech Quality Assessment (SQA) – free form, multi aspect descriptions for a single utterance. 2. Speech Quality Comparison (SQC) – pairwise comparison of two utterances with decision + justification. 3. Speech Quality Improvement Suggestion (SQI) – actionable suggestions to improve a suboptimal utterance. 4. Deepfake Speech Detection (DSD) – classify s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy