OmniAgentBench Dataset Overview OmniAgentBench is a benchmark for evaluating multimodal agents under realistic "wild" conditions: speech input, acoustic noise, dense/scattered instructions, and multi turn conversations. It wraps three existing agent benchmarks (MPCC, GUI Odyssey, EmbodiedBench) with speech audio, noise overlays, and wild text rewrites so that the same tasks can be evaluated under controlled input modality variations. Dataset Structure Benchmarks 1. MPCC (Multi Modal Planning and Control Challenge) Constraint planning over visual schedules (flights, calendars, meetings). Location: mpcc/ , dataset/mpcc/ 9 sub tasks: 3 tasks (flight, calendar, meeting) × 3 difficulties ~2,700 speech samples + ~5,700 task screenshots Metric: Feasible Plan Accuracy (FPA) 2. GUI Odyssey Cross app mobile GUI navigation with speech based instructions. Location: gui odyssey/ 6 app categories, 1,800+ audio samples with screenshots Metric: Action Matching Score (AMS) 3. EmbodiedBench (ALFRED) Vision driven household tasks in AI2 THOR environments. Location: embodiedbench/ 300 samples, clean audio + 7 noise environment variants Metric: LCS Ratio (plan sequence similarity) Wild Conditions Each…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy