OmniAgentBench Dataset Card Overview OmniAgentBench is a comprehensive benchmark for evaluating multimodal agents under realistic "wild" conditions including speech input, acoustic noise, and complex multi turn interactions. Organization: omniagentbench Dataset: OmniAgentBench Total Size: 27.1 GB Contributors: Hodfa71, acbueff Dataset Structure The dataset is organized into separate benchmark folders at the root level: ⚠️ Important: Folder Separation The images/ folder contains ONLY MPCC screenshots! Folder Benchmark Content images/mpcc/ MPCC 5,700 screenshots (flight schedules, calendars, meetings) gui odyssey/screenshots/ GUI Odyssey 1,950 mobile app screenshots gui odyssey/General Tool/screenshots/ GUI Odyssey Per category screenshots They are NOT mixed each benchmark has its own dedicated folder. Benchmarks 1. MPCC (Multi Modal Planning and Control Challenge) Constraint planning over visual schedules (flights, calendars, meetings). Location: mpcc/ , images/mpcc/ , dataset/mpcc/ Contents: Audio: 300 speech samples across 3 tasks × 3 difficulties Images: 5,700 screenshots (schedules, flight info, calendars) Text: Structured task instructions with JSON output format Ground Truth:…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy