WildClawBench Hard, practical, end to end evaluation for AI agents — in the wild. WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end to end, without hand holding? We drop agents into a live OpenClaw environment — the same open source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highlights from a football match, negotiating meeting times over multi round emails, hunting down contradictions in search results, writing inference scripts for undocumented codebases, catching privacy leaks before they happen. Useful things. Hard things. Hard enough that the strongest frontier model we tested still tops out around 62% overall (technical report Main results table), and most models land well below that. That makes scores mean something. Why WildClawBench? Most agent benchmarks test isolated capabilities — calling a function, parsing JSON, following a single instruction. WildClawBench tests the full picture: What We Test Why It's Hard : : 🔗 Agency Multi step tool orchestration, error recovery, autonomous planning Agents must chain 10–60+ tool calls, adapt when services fail, and d…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy