KB Whisper Large The National Library of Sweden releases a new suite of Whisper models trained on over 50,000 hours of Swedish speech. In evaluations across FLEURS, CommonVoice and NST, our best performing model reduces the Word Error Rate (WER) by an average of 47% compared to OpenAI's whisper large v3 . The performance of smaller Whisper model sizes on Swedish speech has also substantially improved, with kb whisper small outperforming openai/whisper large v3 (a model six times its size). Model size FLEURS CommonVoice NST tiny KBLab 13.2 12.9 11.2 OpenAI 59.2 67.8 85.2 base KBLab 9.1 8.7 7.8 OpenAI 39.6 52.1 53.4 small KBLab 7.3 6.4 6.6 OpenAI 20.6 26.4 26.4 medium KBLab 6.6 5.4 5.8 OpenAI 12.1 15.8 17.1 large v3 KBLab 5.4 4.1 5.2 OpenAI 7.8 9.5 11.3 Table: Word Error Rate (WER) comparison between KBLab's Whisper models and the corresponding OpenAI versions. Usage We provide checkpoints in different formats: Hugging Face , whisper.cpp (GGML), onnx , and ctranslate2 (used in faster whisper and WhisperX ). 2025 05 13 Update! The default when loading our models through Hugging Face is Stage 2 . As of May 2025 there exists two Stage 2 versions in addition to the default, namely Subtit…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy