Dataset Card for OpenAI HumanEval Table of Contents OpenAI HumanEval Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Initial Data Collection and Normalization Who are the source language producers? Annotations Annotation process Who are the annotators? Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Repository: GitHub Repository Paper: Evaluating Large Language Models Trained on Code Dataset Summary The HumanEval dataset released by OpenAI includes 164 programming problems with a function sig nature, docstring, body, and several unit tests. They were handwritten to ensure not to be included in the training set of code generation models. Supported Tasks and Leaderboards Languages The programming problems are written in Python and contain English natural text in comments and docstrings. Dataset Structure Data Instances An exampl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy