RoadmapBench A benchmark for evaluating AI coding agents on multi target, long horizon software development tasks derived from open source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Quick Start Prerequisites Docker ( = 20.0) Python = 3.9 Harbor (recommended for agent evaluation) 1. Download Dataset Note : Pre built Docker images are available on DockerHub ( znpt/roadmapbench ). When using Harbor (recommended), images are pulled automatically — no local build required. The Docker images contain the full git history (truncated at V OLD) for agent use, while the HuggingFace dataset provides source code without git history for smaller download size. 2. Run Agent Evaluation (Recommended: Harbor) Harbor provides out of the box support for running AI agents against RoadmapBench tasks. Terminus 2 agent: OpenHands agent: See full documentation for advanced options (thinking parameters, network isolation, metrics computation). T…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy