BrowseComp Plus BrowseComp Plus is a new benchmark for Deep Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep Research agents . The benchmark sources challenging, reasoning intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both human verified evidence documents sufficient to answer the queries, and mined hard negatives to keep the task challenging. For more information, please see our Project Page, Paper on Hugging Face, or Arxiv. Code: GitHub repository How to use this dataset? Evaluating end to end agents: Use BrowseComp Plus to test combinations of retrievers and LLM agents, evaluated based on the generated answer against the ground truth answer. Evaluating retriever only: Beyond ground truth questions and answers, BrowseComp Plus provides human labels for: Evidence documents (needed to answer the query) Gold documents (needed to answer, and contains the final answer) from which we may measure standard retrieval metrics like nDCG@10. Dataset Structure This da…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy