PrimBench v2 — Full Trajectory Results Browser agent benchmark trajectories from the PrimBench v2 sweep (stock browser use harness, post hardening). Code repo : https://github.com/i2x ai lab/PrimBench Results PR : https://github.com/i2x ai lab/PrimBench/pull/2 Task picks : 1049 per model (519 clean + 530 intervention) Harness : stock browseruse eval.py with concurrency=4, max steps=40, timeout=600s Directory layout per model Models in this release (4/6 complete) Model Pass % Avg Score Gemini 3 Flash 52.9% 0.701 GPT 5.4 48.1% 0.657 GPT 5.4 mini 24.2% 0.387 Qwen3 VL 235B 7.9% 0.143 Claude Opus 4.7 and Gemini 3.1 Pro are still running — will be added on completion. The summary.json files in each model dir have the full per task eval; the aggregate numbers above merge summary.json + per task trajectory.json.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy