Nemotron H 4B Base 8K Model Developer: NVIDIA Model Dates: October 2024 March 2025 Data Freshness: September 2024 The pretraining data has a cutoff date of September 2024. Model Overview NVIDIA Nemotron H 4B Base 8K is a large language model (LLM) developed by NVIDIA, designed as a completion model for a given piece of text. It uses a hybrid model architecture that consists primarily of Mamba 2 and MLP layers combined with just four Attention layers. The model is pruned and distilled from Nemotron H 8B Base 8K using 380B tokens, and features an 8K context length. The supported languages include: English, German, Spanish, French, Italian, Korean, Portuguese, Russian, Japanese, and Chinese. For best performance on a given task, users are encouraged to customize the model using the NeMo Framework suite of customization tools, including Parameter Efficient Fine Tuning (P tuning, Adapters, LoRA, and more), and Model Alignment (SFT, SteerLM, RLHF, and more) using NeMo Aligner. The model was pruned and distilled from Nemotron H Base 8K using our hybrid language model compression technique and then fine tuned into Nemotron H 4B Instruct 128K. For more details, please refer to the paper. Th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy