Read our How to Run Nemotron 3 Nano Omni Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. Model Overview Description: NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end to end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This model was improved using Qwen3 VL 30B A3B Instruct, Qwen3.5 122B A10B, Qwen3.5 397B A17B, Qwen2.5 VL 72B Instruct, and gpt oss 120b. For more information, please see the Training Dataset section below. License/Terms of Use Governing Terms: Use of this model is governed by the NVIDIA Open Model Agreement Deployment Geography: Global Use Case: This model is designe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy