Dataset Summary SWE rebench is a large scale dataset designed to support training and evaluation of LLM based software engineering (SWE) agents, building upon and expanding our earlier release, SWE bench extra. It is constructed using a fully automated pipeline that continuously extracts real world interactive SWE tasks from GitHub repositories at scale, as detailed in our paper SWE rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. The dataset currently comprises over 21,000 issue–pull request pairs from 3,400+ Python repositories, each validated for correctness through automated environment setup and test execution. A curated subset of these tasks also forms the basis of our continuously updated SWE rebench leaderboard. SWE rebench builds upon and extends the methodology of SWE bench by incorporating several key enhancements detailed in our paper, including: A fully automated pipeline for continuous task collection. LLM driven extraction and validation of environment installation instructions. An automated LLM based task quality assessment pipeline that annotates tasks with labels such as clarity, complexity, or test p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy