At a Glance Total parameters 31B (Mamba2 Transformer hybrid MoE) Active parameters ~3B per token Max context 256k tokens Modalities (in) Video, Audio, Image, Text Modality (out) Text Reasoning mode On by default; toggle via enable thinking Best for Video+speech analysis, document intelligence (OCR/charts/long docs), GUI/agentic workflows, ASR Minimum GPU (BF16) 1× H100 80GB (single GPU); 1× B200 / 1× H200 recommended Minimum GPU (FP8) 1× L40S 48GB; 1× RTX Pro 6000 / 1× B200 recommended Minimum GPU (NVFP4) 1× RTX 5090 32GB; 1× DGX Spark / 1× Jetson Thor also supported Precisions BF16 (62 GB) · FP8 (33 GB) · NVFP4 (21 GB) Quick Start Guide Model Parameters Mode temperature top p top k max tokens reasoning budget grace period Thinking mode 0.6 0.95 — 20480 16384 1024 Instruct mode 0.2 — 1 1024 — — Model Overview Description: NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (O…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy