Read our How to Run Nemotron 3 Ultra Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. You can view KLD Divergence Benchmarks, in our guide. NVIDIA Nemotron 3 Ultra 550B A55B Model Summary : : Total Parameters 550B (55B active) Architecture LatentMoE Mamba 2 + MoE + Attention hybrid with Multi Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Korean, Brazilian Portuguese, and Chinese Best For Frontier reasoning, complex agentic workflows, long context analysis, tool use, multilingual reasoning, high stakes RAG Reasoning Mode Configurable on/off via chat template ( enable thinking=True/False ) License OpenMDW License Agreement, version 1.1 Release Date June 4, 2026 Quick Start For more details on how to deploy and use the model see the Quick Start Guide below! For running Nemotron 3 Ultra on a smaller footprint, please see: NVIDIA Nemotron 3 Ultra 550B A55B NVFP4 Model Overview Model Developer: NVIDIA Corporation Model Dates: December 2025 April 2026 Data Freshness: The post training data has a cutoff date…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy