Dataset Summary SWE bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post PR behavior as the reference solution. The dataset was released as part of SWE bench: Can Language Models Resolve Real World GitHub Issues? Want to run inference now? This dataset only contains the problem statement (i.e. issue text) and the base commit which can represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets. princeton nlp/SWE bench oracle princeton nlp/SWE bench bm25 13K princeton nlp/SWE bench bm25 27K princeton nlp/SWE bench bm25 40K princeton nlp/SWE bench bm25 50k llama Supported Tasks and Leaderboards SWE bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. D…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy