WebDS: A Benchmark for Web based Data Science WebDS is the first end to end benchmark designed for evaluating agents on real world web based data science workflows. It contains 870 tasks across 29 containerized websites spanning 10 domains , including economics, health, climate, and scientific research. Agents are tested on: Multi hop web navigation Structured and unstructured data processing Tool usage (e.g., Python scripts, visualization tools) Downstream task completion (e.g., reports, Reddit posts) Tasks reflect realistic data science scenarios, such as acquiring data from government portals, comparing datasets across sites, and synthesizing insights in report ready formats. 📦 Contents This repository includes: tasks/ : JSON files for all 870 benchmark tasks, with metadata and intents websites/ : Dockerized replicas of 29 benchmark sites for reproducibility webds experiments/ : Code for running LLM based agents and collecting evaluation metrics 🌍 Hosted Demo (Docker) You can try a live version of the benchmark via: http://ec2 18 220 211 153.us east 2.compute.amazonaws.com:3333 This is useful for previewing the benchmark environment or debugging agent behavior before running l…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy