The Heap Dataset We develop The Heap , a new contamination free multilingual code dataset comprising 57 languages, which facilitates LLM evaluation reproducibility. The reproduction packge can be found here. Is your code in The Heap? If you would like to have your data removed from the dataset, follow the instructions on GitHub. Citation If you use this dataset as part of your research please cite us: Usage Using the Datasets API, our dataset can be used as follows: Collection We collect up to 50,000 public repositories using the GitHub API, focusing on license type , star count , and creation date . Repositories with non permissive licenses are prioritized to reduce contamination, as public code datasets we deduplicate against primarily focus on permissive or no license repositories. We select repositories created before August 2024 in decreasing order of their star counts. To handle GitHub rate limits, we use timeouts and pagination during the scraping process. Copyleft licenses included in the The Heap License Family : : : : CECILL 1.0, CECILL 1.1, CECILL 2.0, CECILL 2.1, CECILL C, EPL 1.0, EPL 2.0, LGPL 2.1, LGPL 3.0, MS RL, MPL 2.0 Weak Copyleft GPL 2.0, GPL 3.0 Strong Copylef…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy