Fine tuned French Voxpopuli v2 wav2vec2 base model for speech to phoneme task in French Fine tuned facebook/wav2vec2 base fr voxpopuli v2 for French speech to phoneme (without language model) using the train and validation splits of Common Voice v13. Audio samplerate for usage When using this model, make sure that your speech input is sampled at 16kHz . Output As this model is specifically trained for a speech to phoneme task, the output is sequence of IPA encoded words, without punctuation. If you don't read the phonetic alphabet fluently, you can use this excellent IPA reader website to convert the transcript back to audio synthetic speech in order to check the quality of the phonetic transcription. Training procedure The model has been finetuned on Commonvoice v13 (FR) for 14 epochs on a 4x2080 Ti GPUs at Cnam/LMSSC using a ddp strategy and gradient accumulation procedure (256 audios per update, corresponding roughly to 25 minutes of speech per update 2k updates per epoch) Learning rate schedule : Double Tri state schedule Warmup from 1e 5 for 7% of total updates Constant at 1e 4 for 28% of total updates Linear decrease to 1e 6 for 36% of total updates Second warmup boost to 3e…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy