Model Overview Minitron 8B Base is a large language model (LLM) obtained by pruning Nemotron 4 15B; specifically, we prune model embedding size, number of attention heads, and MLP intermediate dimension. Following pruning, we perform continued training with distillation using 94 billion tokens to arrive at the final model; we use the continuous pre training data corpus used in Nemotron 4 15B for this purpose. Deriving the Minitron 8B and 4B models from the base 15B model using our approach requires up to 40x fewer training tokens per model compared to training from scratch; this results in compute cost savings of 1.8x for training the full model family (15B, 8B, and 4B). Minitron models exhibit up to a 16% improvement in MMLU scores compared to training from scratch, perform comparably to other community models such as Mistral 7B, Gemma 7B and Llama 3 8B, and outperform state of the art compression techniques from the literature. Please refer to our arXiv paper for more details. This model is for research and development only. Model Developer: NVIDIA Model Dates: Minitron 8B Base was trained between February 2024 and June 2024. License Minitron 8B Base is released under the NVIDIA…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy