APEX v1 extended The AI Productivity Index (APEX) is a benchmark from Mercor for assessing whether frontier models are capable of performing economically valuable tasks across four jobs: investment banking associate , management consultant , big law associate , and primary care physician (MD) . APEX v1 extended doubles the heldout evaluation set from n=200 to n=400, with increased complexity and variety. On average, tasks take over two and a half hours for seasoned professionals to complete. Tasks: 400 heldout (100 per job) + 100 open source dev set Domains: Investment banking, Management consulting, Law, Medicine Grading: Rubric based, using a Judge LM (Gemini 2.5 Pro, Thinking=On) Runs per model: 8 License: CC BY 4.0 Intended use: APEX v1 extended is intended exclusively for model evaluation. Any use of this dataset for training, fine tuning, or parameter fitting is forbidden. Crawling or scraping the dataset is also forbidden. Dataset overview Each case consists of a prompt, source documents, and a grading rubric with prompt specific quality criteria. Cases were created by 76 experts with a mean of 7.25 years of professional experience, sourced through the Mercor platform. Domai…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy