Dataset Summary SWE bench Lite is subset of SWE bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post PR behavior as the reference solution. The dataset was released as part of SWE bench: Can Language Models Resolve Real World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton nlp/SWE bench Lite.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy