Llama 3.1 Nemotron Ultra 253B v1 FP8 Model Overview Llama 3.1 Nemotron Ultra 253B v1 FP8 is a large language model (LLM) which is a derivative of Meta Llama 3.1 405B Instruct (AKA the reference model ). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model supports a context length of 128K tokens. This model fits on a single 8xH100 node for inference. Llama 3.1 Nemotron Ultra 253B v1 FP8 is a model which offers a great tradeoff between model accuracy and efficiency. Efficiency (throughput) directly translates to savings. Using a novel Neural Architecture Search (NAS) approach, we greatly reduce the model’s memory footprint, enabling larger workloads, as well as reducing the number of GPUs required to run the model in a data center environment. This NAS approach enables the selection of a desired point in the accuracy efficiency tradeoff. Furthermore, by using a novel method to vertically compress the model (see details here), it also offers a significant improvement in latency. The model underwent a multi phase post training process to enhance both its reasoning and non reasoning capabilities. This inc…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy