Dataset Card for CodeSearchNet corpus Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://wandb.ai/github/CodeSearchNet/benchmark Repository: https://github.com/github/CodeSearchNet Paper: https://arxiv.org/abs/1909.09436 Data: https://doi.org/10.5281/zenodo.7908468 Leaderboard: https://wandb.ai/github/CodeSearchNet/benchmark/leaderboard Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language modeling : The dataset can be used to train…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy