HuBERT fine tuned on DUSHA dataset for speech emotion recognition in russian language The pre trained model is this one facebook/hubert large ls960 ft The DUSHA dataset used can be found here Fine tuning Fine tuned in Google Colab using Pro account with A100 GPU Freezed all layers exept projector, classifier and all 24 HubertEncoderLayerStableLayerNorm layers Used half of the train dataset Training parameters 2 epochs train batch size = 8 eval batch size = 8 gradient accumulation steps = 4 learning rate = 5e 5 without warm up and decay Metrics Achieved accuracy = 0.86 balanced = 0.76 macro f1 score = 0.81 on test set, improving accucary and f1 score compared to dataset baseline Usage
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy