WideSearch: Benchmarking Agentic Broad Info Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information seeking tasks. Unlike existing benchmarks that focus on finding a single, hard to find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy to find, information. The challenge in these tasks lies not in cognitive difficulty, but in the operational scale, repetitiveness, and the need for Completeness and Factual Fidelity in the final result. For example, a financial analyst gathering key metrics for all companies in a sector, or a job seeker collecting every vacancy that meets their criteria. The benchmark, originating from the research paper "WideSearch: Benchmarking Agentic Broad Info Seeking," contains 200 meticulously designed tasks (100 in English, 100 in Chinese). See our paper and github repo for more details. Dataset Structure The dataset consists of these components: a task file, and a directory containing the ground truth answers. Data Instances widesearch.jsonl is JSON Lines file, where each line represents a s…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy