Mistral NeMo Minitron 8B Base Model Overview Mistral NeMo Minitron 8B Base is a base text to text model that can be adopted for a variety of natural language generation tasks. It is a large language model (LLM) obtained by pruning and distilling the Mistral NeMo 12B; specifically, we prune the embedding dimension and MLP intermediate dimension in the model. Following pruning, we perform continued training with distillation using 380 billion tokens to arrive at the final model; we use the continuous pre training data corpus used in Nemotron 4 15B for this purpose. Please refer to our technical report for more details. Model Developer: NVIDIA Model Dates: Mistral NeMo Minitron 8B Base was trained between July 24, 2024 and August 10, 2024. License This model is released under the NVIDIA Open Model License Agreement. Model Architecture Mistral NeMo Minitron 8B Base uses a model embedding size of 4096, 32 attention heads, MLP intermediate dimension of 11520, with 40 layers in total. Additionally, it uses Grouped Query Attention (GQA) and Rotary Position Embeddings (RoPE). Architecture Type: Transformer Decoder (Auto Regressive Language Model) Network Architecture: Mistral NeMo Input Typ…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy