Mellum2 Thinking [!Note] Use this model when you want explicit chain of thought before the final answer — complex debugging, multi step planning, agentic workflows, and math or reasoning heavy tasks. For direct, low latency answers without reasoning traces, use Instruct instead. Mellum2 Thinking Highlights Mellum 2 Thinking is a post trained reasoning augmented assistant model trained by JetBrains. The model uses a Mixture of Experts architecture with 64 experts and activates 8 experts per token. It uses a combination of sliding window and full attention layers, with a context length of 131,072 tokens. It is produced from Mellum2 12B A2.5B Base by supervised fine tuning (loss computed only on the final assistant turn) followed by reinforcement learning with verifiable rewards (RLVR) on a harder data mix that includes a long form math subset. The model emits its reasoning inside ... blocks before the final answer. Mellum2 Model Family This repository contains one checkpoint from the Mellum 2 family. Checkpoint Description Base Pretrain Base checkpoint before long context extension Base Final base model Instruct SFT Supervised instruction tuned checkpoint Thinking SFT Supervised thin…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy