DeepSearchQA A 900 prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi step information seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post\ ▶ DeepSearchQA Leaderboard on Kaggle\ ▶ Technical Report\ ▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900 prompt benchmark for evaluating agents on difficult multi step information seeking tasks across 17 different fields. Unlike traditional benchmarks that target single answer retrieval or broad spectrum factuality, DeepSearchQA features a dataset of challenging, hand crafted tasks designed to evaluate an agent’s ability to execute complex search plans to generate exhaustive answer lists. Each task is structured as a "causal chain", where discovering information for one step is dependent on the successful completion of the previous one, stressing long horizon planning and context retention. All tasks are grounded in the open web with objectively verifiable answer sets. DeepSearchQA is meant to be used to evaluate LLMs or LLM agents with access to the web. Dataset Description This dataset is a collection of 900 examples. Each example is composed of: A problem ( pr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy