NVIDIA Nemotron 3 Super 120B A12B Base Model Overview Model Developer: NVIDIA Corporation Model Dates: December 2025 January 2026 Data Freshness: The post training data has a cutoff date of February 2026. The pre training data has a cutoff date of December 2025. Description Nemotron 3 Super 120B A12B Base is a base large language model (LLM) trained from scratch by NVIDIA, with next token prediction loss. It provides a good starting platform for further training, including instruction follow, and coding. The model employs a hybrid Latent Mixture of Experts (LatentMoE) architecture, utilizing interleaved Mamba 2 and MoE layers, along with select Attention layers. Distinct from the Nano model, the Super model incorporates Multi Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using NVFP4 quantization to maximize compute efficiency. The model has 12B active parameters and 120B parameters in total . The supported languages include: English, Spanish, French, German, Japanese, Italian, Chinese, Arabic, Hebrew, Hindi, Korean, Czech, Danish, Dutch, Finnish, Polish, Portuguese, Thai, Swedish, and Russian. This model is ready for commercial use…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy