Fine tuned XLSR 53 large model for speech recognition in Polish Fine tuned facebook/wav2vec2 large xlsr 53 on Polish using the train and validation splits of Common Voice 6.1. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine tuned thanks to the GPU credits generously given by the OVHcloud :) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2 sprint Usage The model can be used directly (without a language model) as follows... Using the HuggingSound library: Writing your own inference script: Reference Prediction """CZY DRZWI BYŁY ZAMKNIĘTE?""" PRZY DRZWI BYŁY ZAMKNIĘTE GDZIEŻ TU POWÓD DO WYRZUTÓW? WGDZIEŻ TO POM DO WYRYDÓ """O TEM JEDNAK NIE BYŁO MOWY.""" O TEM JEDNAK NIE BYŁO MOWY LUBIĘ GO. LUBIĄ GO — TO MI NIE POMAGA. TO MNIE NIE POMAGA WCIĄŻ LUDZIE WYSIADAJĄ PRZED ZAMKIEM, Z MIASTA, Z PRAGI. WCIĄŻ LUDZIE WYSIADAJĄ PRZED ZAMKIEM Z MIASTA Z PRAGI ALE ON WCALE INACZEJ NIE MYŚLAŁ. ONY MONITCENIE PONACZUŁA NA MASU A WY, CO TAK STOICIE? A WY CO TAK STOICIE A TEN PRZYRZĄD DO CZEGO SŁUŻY? A TEN PRZYRZĄD DO CZEGO SŁUŻY NA JUTRZEJSZYM KOLOKWIUM BĘDZIE PIĘĆ PYTAŃ OTWARTYCH I TEST WIELOKROTNEGO WYBOR…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy