Every Eval Ever Datastore This is the datastore for the Every Eval Ever project. The readme from the project GitHub is below. It describes how to submit new benchmarks and evals to this dataset. EvalEval Coalition — "We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations." Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused. The three components that make it work: 📋 A metadata schema ( eval.schema.json ) that defines the information needed for meaningful comparison of evaluation results, including instance level data 🔧 Validation that checks data against the schema before it enters the repository 🔌 Converters for Inspect AI, HELM, and lm eval harness, so you can transform your existing evaluation logs into the standard format Flat datastore view The canonical datastore view is being migrated to a flat, manifest indexed layout under flat/ . The…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy