Model for Dimensional Speech Emotion Recognition based on Wav2vec 2.0 Please note that this model is for research purpose only. A commercial license for a model that has been trained on much more data can be acquired with audEERING. The model expects a raw audio signal as input, and outputs predictions for arousal, dominance and valence in a range of approximately 0...1. In addition, it provides the pooled states of the last transformer layer. The model was created by fine tuning Wav2Vec2 Large Robust on MSP Podcast (v1.7). The model was pruned from 24 to 12 transformer layers before fine tuning. An ONNX export of the model is available from doi:10.5281/zenodo.6221127. Further details are given in the associated paper and tutorial. Usage
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy