REPRO Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research? This repository contains the REPRO Bench dataset, introduced in the paper REPRO Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?. REPRO Bench is a novel benchmark designed to evaluate the capability of agentic AI systems in automating the assessment of social science paper reproducibility. It addresses limitations of existing benchmarks by providing a collection of 112 task instances, each representing a social science paper with a publicly available reproduction report. The benchmark features end to end evaluation tasks on reproducibility with complexity comparable to real world assessments, encompassing diverse data formats and programming languages. Official Codebase: https://github.com/uiuc kang lab/REPRO Bench 📦 Dataset Structure The REPRO Bench dataset includes: 112 task instances (original paper PDFs + corresponding code and data) Gold standard reproducibility annotations Public reproduction reports 🚀 Getting Started The REPRO Bench dataset is hosted on Hugging Face. To access and clone the dataset: To run the representative AI agents evaluated in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy