BrowseComp Long Context BrowseComp Long Context is a dataset based on BrowseComp to benchmark LLM’s capability to retrieve relevant information from noisy data in its context. It converts the agentic question answering tasks from Browsecomp into long context tasks. For each of the questions in a subset of BrowseComp, a list of urls are attached. Each url will be paired with an indicator indicating whether the content of the web page is required to answer the question or is additional content served as supplement information or noise. The required urls are collected and reviewed by a human to ensure they are sufficient and necessary to answer the original question. The additional urls are obtained by searching relevant questions that can help answer the original question. The data is extensible to different context windows, with the provided list of urls, it’s feasible to construct model prompts beyond 1m context window. This eval is challenging because: The constructed prompt is based on real data where most of the context is somewhat relevant, as opposed to a broad web corpus where very little data is relevant The model must combine multiple pieces of information in order to answe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy