💻 Code · 🤗 Models & Data · 📜 Paper · 📓 Blog [!NOTE] For full information, go check out the Tmax paper here. TMax 27B TMax 27B is a model trained using DPPO on top of Qwen 3.6 27B for use as a terminal agent. It achieves roughly 43% on Terminal Bench 2.0 after 160 steps of RL training. This model is part of a collection of terminal agents in various sizes. Additionally, we provide model checkpoints as branches of the repository. The main model checkpoint is step 160 as this performed best on TBLite. For this model only, we upload checkpoints at step 100, 160, 200, 240, 300 steps. Evaluation Results Model TB Lite TB 2.1 TB 2.0 (daytona) Qwen 3.5 2B 5.71 +/ 1.6 1.9 +/ 1.4 2.3 +/ 1.0 Tmax 2B 11.8 +/ 1.4 4.2 +/ 1.2 2.9 +/ 0.6 Qwen 3.5 4B 31.8 +/ 3.8 ? 16.6 +/ 1.7 Tmax 4B 42.6 +/ 1.5 19.9 +/ 1.1 18.9 +/ 1.9 Qwen 3.5 9B 41.9 +/ 2.7 16.1 +/ 3.7 21.1 +/ 2.6 Tmax 9B (this model!) 57.2 +/ 2.5 28.8 +/ 3.7 27.2 +/ 1.5 Qwen 3.6 27B 70.8 +/ 2.1 40.5 +/ 2.4 39.6 +/ 2.1 Tmax 27B 68.6 +/ 4.7 44.9 +/ 1.8 42.7 +/ 0.7 For details on evaluation methodology please check our paper. In general, we used a podman (docker) backend with default timeouts and custom harness similar to mini swe agent. For…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy