WavLM Large for Voice (Sounding) Quality Classification Model Description This model includes the implementation of voice quality classification described in Vox Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits (https://arxiv.org/pdf/2505.14648) Metric: Specifically, we report speaker level Macro F1 scores. Specifically, we randomly sampled five utterances for each speaker and repeated this stratification process 20 times. The speaker level score is computed as the average Macro F1 across speakers. We then report the unweighted average of speaker level Macro F1 scores between VoxCeleb and Expresso. Special Note: We exclude EARS from ParaSpeechCaps due to its limited number of samples in the holdout set. The included labels are: [ 'shrill', 'nasal', 'deep', Pitch 'silky', 'husky', 'raspy', 'guttural', 'vocal fry', Texture 'booming', 'authoritative', 'loud', 'hushed', 'soft', Volume 'crisp', 'slurred', 'lisp', 'stammering', Clarity 'singsong', 'pitchy', 'flowing', 'monotone', 'staccato', 'punctuated', 'enunciated', 'hesitant', Rhythm ] Library: https://github.com/tiantiaf0627/vox profile release How to use this model Download repo Inst…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy