Mistral NeMo Minitron 8B Instruct Model Overview Mistral NeMo Minitron 8B Instruct is a model for generating responses for various text generation tasks including roleplaying, retrieval augmented generation, and function calling. It is a fine tuned version of nvidia/Mistral NeMo Minitron 8B Base, which was pruned and distilled from Mistral NeMo 12B using our LLM compression technique. The model was trained using a multi stage SFT and preference based alignment technique with NeMo Aligner. For details on the alignment technique, please refer to the Nemotron 4 340B Technical Report. The model supports a context length of 8,192 tokens. Try this model on build.nvidia.com. Model Developer: NVIDIA Model Dates: Mistral NeMo Minitron 8B Instruct was trained between August 2024 and September 2024. License NVIDIA Open Model License Model Architecture Mistral NeMo Minitron 8B Instruct uses a model embedding size of 4096, 32 attention heads, MLP intermediate dimension of 11520, with 40 layers in total. Additionally, it uses Grouped Query Attention (GQA) and Rotary Position Embeddings (RoPE). Architecture Type: Transformer Decoder (Auto regressive Language Model) Network Architecture: Mistral N…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy