Whisper Large V3 for Categorical Emotion Classification Model Description This model includes the implementation of categorical emotion classification described in Vox Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits (https://arxiv.org/pdf/2505.14648) The training pipeline used is also the top performing solution (SAILER) in INTERSPEECH 2025—Speech Emotion Challenge (https://lab msp.com/MSP Podcast Competition/IS2025/). Note that we did not use all the augmentation and did not use the transcript compared to our official challenge submission system, but we created a speech only system to make the model simple but still effective. We use the MSP Podcast data to train this model, noting that the model might be sensitive to content information when making emotion predictions. However, this could be a good feature for classifying emotions from online content. The included emotions are: [ 'Anger', 'Contempt', 'Disgust', 'Fear', 'Happiness', 'Neutral', 'Sadness', 'Surprise', 'Other' ] Library: https://github.com/tiantiaf0627/vox profile release How to use this model Download repo Install the package Load the model Prediction If you have any…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy