All rights and obligations of the dataset are with original authors of the paper/dataset. I have merely made this dataset with a MIT licence available on HuggingFace. BIG Bench Hard Dataset This repository contains a copy of the BIG Bench Hard dataset. Small edits to the formatting of the dataset are made to integrate it into the Inspect Evals repository, a community contributed LLM evaulations for Inspect AI a framework by the UK AI Safety Institute. The BIG Bench Hard dataset is a collection of various task categories, with each task focused on testing specific reasoning, logic, or language abilities. The dataset also includes two types of 3 shot prompts for each task: answer only prompts and chain of thought prompts. Dataset Structure Main Task Datasets The collection includes a wide range of tasks, with each designed to evaluate different aspects of logical reasoning, understanding, and problem solving abilities. Below is a list of all included tasks: 1. Boolean Expressions Evaluate the truth value of a Boolean expression using Boolean constants ( True , False ) and basic operators ( and , or , not ). 2. Causal Judgment Given a short story, determine the likely answer to a caus…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy