Granite speech 3.2 8b Model Summary: Granite speech 3.2 8b is a compact and efficient speech language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite speech 3.2 8b uses a two pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite speech 3.2 8b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated. The model was trained on a collection of public corpora comprising diverse datasets for ASR and AST as well as synthetic datasets tailored to support the speech translation task. Granite speech 3.2 was trained by modality aligning granite 3.2 8b instruct (https://huggingface.co/ibm granite/granite 3.2 8b instruct) to speech on publicly available open source corpora containing audio inputs and text targets. Evaluations: We evaluated granite speech 3.2 8b alongside other speech language models (SLMs) in the less than 8b parameter range as well as dedicated ASR and AST systems on standard benchmarks. The evaluation spanned multiple p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy