Llama 3.1 Nemotron Nano 8B v1 Model Overview Llama 3.1 Nemotron Nano 8B v1 is a large language model (LLM) which is a derivative of Meta Llama 3.1 8B Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. Llama 3.1 Nemotron Nano 8B v1 is a model which offers a great tradeoff between model accuracy and efficiency. It is created from Llama 3.1 8B Instruct and offers improvements in model accuracy. The model fits on a single RTX GPU and can be used locally. The model supports a context length of 128K. This model underwent a multi phase post training process to enhance both its reasoning and non reasoning capabilities. This includes a supervised fine tuning stage for Math, Code, Reasoning, and Tool Calling as well as multiple reinforcement learning (RL) stages using REINFORCE (RLOO) and Online Reward aware Preference Optimization (RPO) algorithms for both chat and instruction following. The final model checkpoint is obtained after merging the final SFT and Online RPO checkpoints. Improved using Qwen. This model is part of the Llama Nemotron Collection. You can find the other model(…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy