Granite Speech 4.1 2B Plus Model Summary Granite Speech 4.1 2B Plus has similar capabilities to the Granite Speech 4.1 2B model. The plus model adds two new community requested rich transcription features that can be activated with a simple prompt change: speaker attributed ASR (speaker labels and word transcripts) and word level timing information. Unlike the base mode, the plus model doesn't provide punctuation and capitalization. The model was trained on corpora similar to the Granite Speech 4.1 2B model which were augmented with speaker turns and word level timestamp tags. This allows the model to provide different modes of functionality controlled by different prompts. Two additional model variants explore different capabilities and inference optimization: Granite Speech 4.1 2B for applications where accuracy is the primary concern with support for punctuated, capitalized transcripts, AST and keyword biased recognition, and includes Japanese. Granite Speech 4.1 2B NAR introduces a novel non autoregressive architecture for higher throughput ASR only mode In this mode the model generates only the text transcript similar to the Granite Speech 4.1 2B model. Speaker attributed ASR…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy