SWE ZERO 12M Trajectories The largest agentic coding trace dataset to date: 112 B tokens of execution free agentic trajectories covering 122 K pull requests , 3 K repositories , and 16 programming languages . Motivation Agentic mid training has become a standard ingredient for frontier coding models: Code World Model (Sept 2025) mid trains on 5 T tokens, using 3 M agentic trajectories from 10.2 K Docker images across 3.15 K repos. Kimi Dev (Dec 2025) mid trains on ~150 B tokens of "high quality and real world data" to instill BugFixer and TestWriter priors. DeepSeek V4 (May 2026) incorporates agentic data during mid training to enhance coding capability. All three rely on containerized execution to verify trajectories, and that is where the data ceiling sits. Kimi Dev was first to call it out: "while repository snapshots from GitHub are available, not all snapshots are equipped with an executable Docker environment." The current largest containerized collection, SWE rebench V2 , ships 32 K tasks with pre built images, and concedes the ceiling by additionally releasing 120 K+ tasks that could not be containerized (4× more) with only install instructions and fail to pass metadata, be…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy