O11y Bench Leaderboard Submissions This repository accepts leaderboard submissions for o11y bench, Grafana's benchmark for LLM agents on observability and SRE tasks. How to Submit 1. Fork this repository 2. Create a new branch for your submission 3. Add your submission under submissions/o11y bench/1.0/ / 4. Open a Pull Request Submission Structure Required: metadata.yaml Each submission must include a metadata.yaml file with the following fields: Job Directory Requirements Each job directory must contain all of the contents of the Harbor run you want scored, including: config.json result.json all trial subdirectories agent logs and artifacts downloaded with the run verifier output for each trial Validation Rules Submissions are expected to preserve the benchmark's shipped evaluation settings. In the current local o11y bench repo: Harbor runs use timeout multiplier = 1.0 ; submissions should not override it Every task currently ships with agent.timeout sec = 600.0 Every task currently ships with verifier.timeout sec = 300.0 Every task currently ships with environment.build timeout sec = 600.0 No resource overrides ( override cpus , override memory mb , override storage mb ) All tria…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy