Breeze ASR 25 GitHub Paper Breeze ASR 25 是一款基於 Whisper large v2 開發的語音辨識模型,並具有以下特色: 強化繁體中文情境辨識能力 強化中英混用情境辨識能力,包含句內以及句外轉換 強化時間戳記對齊,適合自動字幕生成 Breeze ASR 25 is an advanced ASR model fine tuned from Whisper large v2 Optimized for Taiwanese Mandarin Optimized for Mandarin English code switching scenarios, including intra sentential switching and inter sentential switching. Enhanced time alignment, suitable for automatic captioning Example: 增強範例 中英混用情境: MediaTek's 24th Anniversary Breeze ASR 25: Whisper large v2: Performance Word error rates of benchmarks. The WERR is reported in comparison with the Whisper large v2 automatic language detection (WLV2 Auto) baseline. "Breeze ASR 25" is refered in the paper as "Twister" Short form Audio Datasets Dataset\Model Language WLV2 Auto ↓ WLV3 Auto ↓ COOL Whisper ↓ Breeze ASR 25 (Ours) ↓ ASCEND OVERALL Mixed 21.14 23.22 19.71 17.74 ( 16.08%) ASCEND EN English 27.36 27.21 29.39 26.64 ( 2.63%) ASCEND ZH Mandarin 17.49 17.41 18.90 16.04 ( 8.29%) ASCEND MIX Mixed 21.01 25.13 17.34 16.38 ( 22.01%) CommonVoice16 zh TW Mandarin 9.84 8.95 11.86 7.97 ( 19%) CSZS zh en Mixed 29.49 26.43 20.90 13.01 ( 55.88%) Long form Audio Datasets Dataset\Model Language WLV2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy