Model for Age and Gender Recognition based on Wav2vec 2.0 (6 layers) The model expects a raw audio signal as input and outputs predictions for age in a range of approximately 0...1 (0...100 years) and gender expressing the probababilty for being child, female, or male. In addition, it also provides the pooled states of the last transformer layer. The model was created by fine tuning Wav2Vec2 Large Robust on aGender, Mozilla Common Voice, Timit and Voxceleb 2. For this version of the model we only trained the first six transformer layers. An ONNX export of the model is available from doi:10.5281/zenodo.7761387. Further details are given in the associated paper and tutorial. Usage
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy