Higgs Audio v3 STT A speech to text model combining a Whisper Large v3 encoder with a Qwen3 decoder (2.68B total parameters). Update (June 2026) This repository now hosts an updated checkpoint. Changes: Fine tuning data refreshed: public train splits of AMI (IHM), VoxPopuli (en), SPGISpeech, LibriSpeech, TED LIUM, GigaSpeech, plus the public Earnings22 train split ( sanchit gandhi/earnings22 split ) with all rows from source recordings that appear in the ESB/Open ASR test sets excluded. transcribe.py adds a phrase level repetition loop collapse alongside the existing word repetition cap (implemented inline in transcribe.py ; ngram loop fix.py carries the standalone reference and tests). Both are deterministic and applied uniformly to every dataset. Evaluation: see the Open ASR Leaderboard for independently produced results; note that figures listed there predate this update until the entry is re evaluated. Figures previously listed on this card came from an earlier checkpoint and evaluation setup and are superseded. The previous weights remain available via the git revision history. Usage Important: This model uses a custom architecture. You must pass trust remote code=True when lo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy