ARES Bench ARES Bench is the open audit substrate released with the paper Auditing LLM User Simulators for Recommender A/B Testing (NeurIPS 2026, ED Track, under review). It turns the ARES reliability audit view — the LLM backbone is the measurement instrument under test , not an interchangeable implementation detail — into a reproducible protocol over structured behavioral logs, a portable visual sandbox, and a screenshot cache. This release hosts the 17,000 session core corpus that underpins the 200 stratified user main experiment reported in the paper. The companion code, analysis toolkit, and sample logs live at the anonymized GitHub repository (MIT licensed). Dataset summary Dimension Value Total behavioral sessions 17,000 Text sandbox sessions 9,000 (9 backbones × 5 recommenders × 200 users) Visual sandbox sessions 8,000 (8 vision capable backbones × 5 recommenders × 200 users) LLM backbones audited 9 (GPT 4.1, GPT 5.1, Claude Sonnet 4, Claude Opus 4.6, Gemini 2.5 Flash, Gemini 3 Flash, DeepSeek V3.2, Qwen3.5 Large (397b a17b), Qwen3.5 Small (35b a3b)) Recommender models 5 (FM, DeepFM, Pop, PrefAlign, Random) Underlying users 200 stratified MovieLens 1M users per cell Interac…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy