Wav2Vec2 BERT Alvin This model is a fine tuned version of facebook/w2v bert 2.0. This has a CER of 10.27 on Common Voice 16 (yue) test set (without punctuations). Training and evaluation data For training, three datasets were used: Common Voice 16 zh HK and yue Train Set CantoMap: Winterstein, Grégoire, Tang, Carmen and Lai, Regine (2020) "CantoMap: a Hong Kong Cantonese MapTask Corpus", in Proceedings of The 12th Language Resources and Evaluation Conference, Marseille: European Language Resources Association, p. 2899 2906. Cantonse ASR: Yu, Tiezheng, Frieske, Rita, Xu, Peng, Cahyawijaya, Samuel, Yiu, Cheuk Tung, Lovenia, Holy, Dai, Wenliang, Barezi, Elham, Chen, Qifeng, Ma, Xiaojuan, Shi, Bertram, Fung, Pascale (2022) "Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset", 2022. Link: https://arxiv.org/pdf/2201.02419.pdf Code Example or Training Hyperparameters learning rate: 5e 5 train batch size: 4 (on 1 3090) eval batch size: 1 gradient accumulation steps: 32 total train batch size: 32x4=128 optimizer: Adam with betas=(0.9,0.999) and epsilon=1e 08 lr scheduler warmup steps: 1500
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy