SWE smith Dataset Code • Paper • Site [12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the SWE bench/SWE smith [lang] datasets. For better maintainability and ease of use, we are maintaining language specific datasets in lieu of this mono repo. The SWE smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE smith toolkit. It is the largest dataset to date for training software engineering agents. All SWE smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy