Granite Speech 4.1 2B NAR Model Summary: Granite Speech 4.1 2B NAR is a non autoregressive (NAR) speech recognition model that formulates ASR as conditional transcript editing. Instead of decoding tokens one at a time, it edits a CTC hypothesis in a single forward pass using a bidirectional LLM, achieving competitive accuracy with faster inference than autoregressive alternatives. The model is based on the NLE (Non autoregressive LLM based Editing) architecture described in this paper. For applications where accuracy is the primary concern, consider granite speech 4.1 2b , an autoregressive model from the Granite Speech 4.1 family which achieves higher transcription accuracy at the cost of increased inference latency. Granite speech 4.1 2b produces punctuated and capitalized transcripts, supports AST and keyword biased recognition, and includes Japanese. When speaker or word timing information is needed, consider using granite speech 4.1 2b plus , which extends the above model with speaker attributed ASR (speaker labels + word transcripts) and word level timing information. Release Date : April 2026 License: Apache 2.0 Supported Languages: English, French, German, Spanish, Portugue…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy