💻 Code · 🤗 Models & Data · 📜 Paper · 📓 Blog [!NOTE] For full information, go check out the Tmax paper here. TMax 9B TMax 9B is a model trained using DPPO on top of Qwen 3.5 9B for use as a terminal agent. It achieves roughly 27% on Terminal Bench 2.0 after 200 steps of RL training, improving by ~6 points over Qwen 3.5 9B with our harness. This model is part of a collection of terminal agents in various sizes. Additionally, we provide model checkpoints as branches of the repository. The main model checkpoint is step 200 as this performed best on TBLite. We also provide all model rollouts and logprobs over the course of training here! This allows you to inspect what went on during our runs. See the folder for more details (I recommend pointing an agent at it to make plots, etc). Evaluation Results Model TB Lite TB 2.1 TB 2.0 (daytona) Qwen 3.5 2B 5.71 +/ 1.6 1.9 +/ 1.4 2.3 +/ 1.0 Tmax 2B 11.8 +/ 1.4 4.2 +/ 1.2 2.9 +/ 0.6 Qwen 3.5 4B 31.8 +/ 3.8 ? 16.6 +/ 1.7 Tmax 4B 42.6 +/ 1.5 19.9 +/ 1.1 18.9 +/ 1.9 Qwen 3.5 9B 41.9 +/ 2.7 16.1 +/ 3.7 21.1 +/ 2.6 Tmax 9B (this model!) 57.2 +/ 2.5 28.8 +/ 3.7 27.2 +/ 1.5 Qwen 3.6 27B 70.8 +/ 2.1 40.5 +/ 2.4 39.6 +/ 2.1 Tmax 27B 68.6 +/ 4.7 4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy