qwen2.5 7b base retool slime sft v2 ReTool SFT cold start of Qwen/Qwen2.5 7B (base, not Instruct), trained with slime on a Megatron backend. This is the tool format SFT stage that teaches the base model the ReTool interleaved code/tool call format before RL. Training configuration Setting Value Base model Qwen/Qwen2.5 7B (base) Dataset ReTool SFT ( messages format) Loss sft loss , per token, Qwen loss mask Global batch size 32 Epochs 6 (~371 optimizer steps over the ~2k row corpus, ~62 steps/epoch) Optimizer Adam (β1=0.9, β2=0.95), weight decay 0.01 LR schedule 1e 5 → 1e 6, cosine, 10% warmup Precision bf16, grads all reduced in fp32 Parallelism TP=4, PP=1, CP=1, sequence parallel Hardware 8× A100 40GB Final train loss ~0.02 The checkpoint corresponds to iteration 371 (end of epoch 6). Usage Notes This is an intermediate SFT cold start checkpoint intended as the starting point for subsequent RL (GRPO / GFlow RL) in the ReTool pipeline.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy