Nemotron Mini 4B Instruct Model Overview Nemotron Mini 4B Instruct is a model for generating responses for roleplaying, retrieval augmented generation, and function calling. It is a small language model (SLM) optimized through distillation, pruning and quantization for speed and on device deployment. It is a fine tuned version of nvidia/Minitron 4B Base, which was pruned and distilled from Nemotron 4 15B using our LLM compression technique. This instruct model is optimized for roleplay, RAG QA, and function calling in English. It supports a context length of 4,096 tokens. This model is ready for commercial use. Try this model on build.nvidia.com. For more details about how this model is used for NVIDIA ACE, please refer to this blog post and this demo video, which showcases how the model can be integrated into a video game. You can download the model checkpoint for NVIDIA AI Inference Manager (AIM) SDK from here. Model Developer: NVIDIA Model Dates: Nemotron Mini 4B Instruct was trained between February 2024 and Aug 2024. License NVIDIA Community Model License Model Architecture Nemotron Mini 4B Instruct uses a model embedding size of 3072, 32 attention heads, and an MLP intermedia…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy