Dataset Summary SWE bench Verified is a subset of 500 samples from the SWE bench test set, which have been human validated for quality. SWE bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human validation process. The dataset collects 500 test Issue Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post PR behavior as the reference solution. The original SWE bench dataset was released as part of SWE bench: Can Language Models Resolve Real World GitHub Issues? Want to run inference now? This dataset only contains the problem statement (i.e. issue text) and the base commit which represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets. princeton nlp/SWE bench Lite oracle princeton nlp/SWE bench Lite bm25 13K princeton nlp/SWE bench Lite bm25 27K Supported Tasks and Leaderboards SWE bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy