Dataset Summary This dataset contains 80,036 trajectories generated by a software engineering agent based on the SWE agent framework, using various models as action generators. In these trajectories, the agent attempts to solve GitHub issues from the nebius/SWE bench extra and the dev split of princeton nlp/SWE bench. Dataset Description This dataset was created as part of a research project focused on developing a software engineering agent using open weight models, which achieved a score of 40.6% on the SWE bench Verified benchmark. The detailed process of achieving this is outlined in our blog post "Leveraging training and search for better software engineering agents". The dataset collection consisted of two stages: collecting issue resolution instances, following a methodology similar to SWE bench, and generating a large number of trajectories for solving the collected issues. The generated code patches in these trajectories were evaluated by the tests from the linked pull requests to determine which of them passed the tests. The detailed process of collecting issue resolution instances is described in our blog post "Scaling Data Collection for Training Software Engineering Ag…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy