Mellum2 Base [!Note] Use this checkpoint as the starting point for your own fine tuning, alignment, or domain adaptation on top of the long context base. For instruction following or reasoning tasks out of the box, use Instruct or Thinking instead. Mellum2 Base Highlights Mellum2 Base is a long context pretrained causal language model trained by JetBrains. The model uses a Mixture of Experts architecture with 64 experts and activates 8 experts per token. It uses a combination of sliding window and full attention layers, with a context length of 131,072 tokens. This is the long context base, produced from Mellum2 12B A2.5B Base Pretrain by a layer selective YaRN extension stage that re maps RoPE frequencies on the global attention layers only. It is the shared starting point for the released Instruct and Thinking variants. Mellum2 Model Family This repository contains one checkpoint from the Mellum2 family. Checkpoint Description Base Pretrain Base checkpoint before long context extension Base Final base model Instruct SFT Supervised instruction tuned checkpoint Thinking SFT Supervised thinking checkpoint Instruct RL tuned instruction model Thinking RL tuned thinking model Model Ove…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy