NVIDIA Nemotron 3 Super 120B A12B BF16 Model Summary : : Total Parameters 120B (12B active) Architecture LatentMoE Mamba 2 + MoE + Attention hybrid with Multi Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100 80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Chinese Best For Agentic workflows, long context reasoning, high volume workloads (e.g. IT ticket automation), tool use, RAG Reasoning Mode Configurable on/off via chat template ( enable thinking=True/False ) License NVIDIA Nemotron Open Model License Release Date March 11, 2026 Quick Start Use temperature=1.0 and top p=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike. For more details on how to deploy and use the model see the Quick Start Guide below! For running Nemotron 3 Super on a single B200 or DGX Spark please see: NVIDIA Nemotron 3 Super 120B A12B NVFP4 Model Overview Model Developer: NVIDIA Corporation Model Dates: December 2025 March 2026 Data Freshness: The post training data has a cutoff date of February 2026. The pre training data has a cutoff date of June 2025. What is Nemotron? NVIDIA Nemotron™ is a family…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy