CodeSearchNetRetrieval An MTEB dataset Massive Text Embedding Benchmark The dataset is a collection of code snippets and their corresponding natural language queries. The task is to retrieve the most relevant code snippet for a given query. Task category t2t Domains Programming, Written Reference https://huggingface.co/datasets/code search net/ Source datasets: code search net/code search net How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a part of the MMTEB Contribution. Dataset Statistics Dataset Statistics The following code contains the descriptive statistics from the task. These can also be obtained using: This dataset card was automatically generated using MTEB
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy