Introduction Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor intensive filtering funnel from billions of queries, these "needles in the haystack" are non factoid, multi perspective questions that probably require a lot of sub questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions. The train dataset has about 90k samples. Use Cases We provide the dataset as is without any code or specific evaluation criteria. For retrieval augmented generation (RAG), the intent would to at least use the content of the clicked documents in the DocStream to ground an LLM's response to the question. Alternatively, you can issue the queries in the queries field to a search engine api and use the retrieved documents for grounding. In both cases, the intended evaluation would be a side by side LLM as a judge to compare your candidate output to e.g. a closed book reference output from GPT 4. This is an open project we invite the community to take on. For ranking/retrieval evaluation, ideally, you would have access to the Clueweb22 corpus and retrieve from…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy