Llama 3.3 Nemotron Super 49B v1.5 FP8 Model Overview Llama 3.3 Nemotron Super 49B v1.5 FP8 is a significantly upgraded version of Llama 3.3 Nemotron Super 49B v1 and is a large language model (LLM) which is a derivative of Meta Llama 3.3 70B Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and agentic tasks, such as RAG and tool calling. The model supports a context length of 128K tokens. Llama 3.3 Nemotron Super 49B v1.5 FP8 is a model which offers a great tradeoff between model accuracy and efficiency. Efficiency (throughput) directly translates to savings. Using a novel Neural Architecture Search (NAS) approach, we greatly reduce the model’s memory footprint, enabling larger workloads, as well as fitting the model on a single GPU at high workloads (H200). This NAS approach enables the selection of a desired point in the accuracy efficiency tradeoff. For more information on the NAS approach, please refer to this paper The model underwent a multi phase post training process to enhance both its reasoning and non reasoning capabilities. This includes a supervised fine tuning stage for Math, Code, Science, and Too…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy