Wav2Vec2 Large Robust finetuned on the revised ETRI Data of Korean English for Pronunciation Model This repository contains a fine tuned Wav2Vec2 Large Robust model for phoneme recognition tasks. The model was trained and evaluated on our in house English pronunciations of Korean learners dataset, which was made with ETRI and revised by SNU SLP lab. Creator & Uploader: Sehyun Oh (ohsehyun12@snu.ac.kr) Data Information Dataset Name : English Pronunciation of Korean Learners (made with ETRI) revised by SNU SLP lab Data Type : Speech recordings of Korean learners speaking English, annotated with phoneme sequences. Annotation : Each utterance is transcribed at the phoneme level, including pronunciation errors marked with err. These errors highlight phoneme substitutions, insertions, and deletions that occur due to the influence of the Korean language on English pronunciation. Train Set : 14,305 samples Valid Set : 1,590 samples Test Set : 3,974 samples Training Procedure The model was fine tuned for phoneme recognition using the Hugging Face transformers library. Below are the training steps: 1. Data preprocessing to align audio with phoneme labels. 2. Wav2Vec2 Large Robust model fine…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy