wav2vec2 large xlsr 53 th Finetuning wav2vec2 large xlsr 53 on Thai Common Voice 7.0 Read more on our blog We finetune wav2vec2 large xlsr 53 based on Fine tuning Wav2Vec2 for English ASR using Thai examples of Common Voice Corpus 7.0. The notebooks and scripts can be found in vistec ai/wav2vec2 large xlsr 53 th. The pretrained model and processor can be found at airesearch/wav2vec2 large xlsr 53 th. robust speech event Add syllable tokenize , word tokenize (PyThaiNLP) and deepcut tokenizers to eval.py from robust speech event Eval results on Common Voice 7 "test": WER PyThaiNLP 2.3.1 WER deepcut SER CER Only Tokenization 0.9524% 2.5316% 1.2346% 0.1623% Cleaning rules and Tokenization TBD TBD TBD TBD Usage Datasets Common Voice Corpus 7.0](https://commonvoice.mozilla.org/en/datasets) contains 133 validated hours of Thai (255 total hours) at 5GB. We pre tokenize with pythainlp.tokenize.word tokenize . We preprocess the dataset using cleaning rules described in notebooks/cv preprocess.ipynb by @tann9949. We then deduplicate and split as described in ekapolc/Thai commonvoice split in order to 1) avoid data leakage due to random splits after cleaning in Common Voice Corpus 7.0 and 2) p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy