Nemotron 3.5 ASR ONNX h1, h2, h3, h4, h5, h6 { color: 76b900; / NVIDIA green / font weight: 700; } hr { border: none; border top: 1px solid e5e7eb; margin: 2rem 0; } / Improve list spacing / ul, ol { margin top: 0.5rem; margin bottom: 0.5rem; } / Badge alignment consistency / img { display: inline; vertical align: middle; } [!Note] This model is the quantized INT4 ONNX version of nvidia/nemotron 3.5 asr streaming 0.6b, adding language ID prompt conditioning to support transcription across 40 language locales from a single model. It supports streaming inference with 0.56 seconds of latency and is simshipped alongside the baseline NVIDIA model. Nemotron 3.5 ASR is a multilingual, streaming Automatic Speech Recognition (ASR) model engineered to deliver high quality multilingual transcription across both low latency streaming and high throughput batch workloads. Developed by NVIDIA, this 600M parameter model transcribes speech into text with native support for punctuation and capitalization, and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 320ms, 560ms, and 1120ms. This ONNX model was exported with optimization for the 560ms chunk size. By leveraging…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy