rubaiSTT 2v Medium Uzbek Speech to Text Model Classic Whisper medium model fine tuned for Uzbek language. The dataset included of diverse audio: publicly available podcasts, Tashkent dialect podcasts, news, google fleurs, USC and Common Voice 17. Data quality was mixed with 50% human transcribed and 50% pseudo transcribed using Gemini 2.5 Pro. Difference between v1 is that v2 is fully open sourced. Due to some conflicts with data partners, v1 was removed, and the 500 hour dataset was excluded. Instead, new and different datasets were included—all of which will be open sourced. Training scripts will also be open sourced. The entire process will be fully repeatable. Special attention was given to Tashkent dialect audio materials, resulting in strong performance on this dialect. Future versions will include other regional dialects to improve overall coverage. Whitepaper For more details on the methodology and research behind this model, visit: https://uz speech.web.app/rubaistt02m Training and filtering code: https://github.com/Islomov49/rubaistt v2 open sourced Support my works and open source movement: https://tirikchilik.uz/islomovs Model Details Base Model: Whisper Medium Paramete…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy