Dataset Card for IFEval Dataset Description Repository: https://github.com/google research/google research/tree/master/instruction following eval Paper: https://huggingface.co/papers/2311.07911 Leaderboard: https://huggingface.co/spaces/open llm leaderboard/open llm leaderboard Point of Contact: Le Hou Dataset Summary This dataset contains the prompts used in the Instruction Following Eval (IFEval) benchmark for large language models. It contains around 500 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times" which can be verified by heuristics. To load the dataset, run: Supported Tasks and Leaderboards The IFEval dataset is designed for evaluating chat or instruction fine tuned language models and is one of the core benchmarks used in the Open LLM Leaderboard. Languages The data in IFEval are in English (BCP 47 en). Dataset Structure Data Instances An example of the train split looks as follows: Data Fields The data fields are as follows: key : A unique ID for the prompt. prompt : Describes the task the model should perform. instruction id list : An array of verifiable instructions. See Table 1 of the paper for the full…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy