π The paper of WebWalkerQA is available at arXiv. π The dataset resource is a collection of 680 questions and answers from the WebWebWalker dataset. π The dataset is in the form of a JSON file. The keys in the JSON include: Question, Answer, Root Url, and Info. The Info field contains more detailed information, including Hop, Domain, Language, Difficulty Level, Source Website, and Golden Path. ποΈ We also release a collection of 15k silver dataset, which although not yet carefully human verified, can serve as supplementary \textbf{training data} to enhance agent performance. π If you have any questions, please feel free to contact us via the Github issue. βοΈ Due to the web changes quickly, the dataset may contain outdated information, such as golden path or source website. We encourage you to contribute to the dataset by submitting a pull request to the WebWalkerQA or contacting us. π‘ If you find this dataset useful, please consider citing our paper:
Runs entirely in your browser via DuckDB-Wasm β this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy