WavLM Large for Categorical Emotion Classification Model Description This model includes the implementation of categorical emotion classification described in Vox Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits (https://arxiv.org/pdf/2505.14648) The training pipeline used is also the top performing solution (SAILER) in INTERSPEECH 2025 Speech Emotion Challenge (https://lab msp.com/MSP Podcast Competition/IS2025/). Note that we did not use all the augmentation and and did not use the transcript to make the model simple but still effective compared to our INTERSPEECH Challenge solution. We use the MSP Podcast data for training this model. The included emotions are: [ 'Anger', 'Contempt', 'Disgust', 'Fear', 'Happiness', 'Neutral', 'Sadness', 'Surprise', 'Other' ] Library: https://github.com/tiantiaf0627/vox profile release How to use this model Download repo Install the package Load the model Prediction If you have any questions, please contact: Tiantian Feng (tiantiaf@usc.edu) Kindly cite our paper if you are using our model or find it useful in your work Responsible use of the Model: the Model is released under Open RAIL license, and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy