Dataset Card for No Robots 🙅♂️🤖 Look Ma, an instruction dataset that wasn't generated by GPTs! Dataset Description Repository: https://github.com/huggingface/alignment handbook Paper: Leaderboard: https://huggingface.co/spaces/HuggingFaceH4/open llm leaderboard Point of Contact: Lewis Tunstall Dataset Summary No Robots is a high quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single turn instructions across the following categories: Category Count : : Generation 4560 Open QA 1240 Brainstorm 1120 Chat 850 Rewrite 660 Summarize 420 Coding 350 Classify 350 Closed QA 260 Extract 190 Supported Tasks and Leaderboards The No Robots dataset designed for instruction fine tuning pretrained language models and we recommend benchmarking against the following: MT Bench: a multi turn benchmark spanning 80 dialogues and 10 domains. AlpacaEval: a single turn benchmark which evaluates the performance of chat and instruct mode…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy