data agent rl environment train The official verified training suite for the data agent RL pipeline. 2238 Harbor format data analysis tasks, each with: An LLM assigned difficulty label (L1 L5) A Kaggle dataset dependency (fetched at container start) A tested reward function This is the training data counterpart to AdithyaSK/data agent rl environment eval . For your held out eval split, use that one. 💡 Browse in your browser — click the badge above or open AdithyaSK/harbor visualiser to inspect every task's spec, instruction, environment, tests, and difficulty. Why "training" vs "eval" This dataset ( train ) Eval ( eval ) Pipeline run Stage 1 only (Sonnet anchor + categorize on pass) Stage 1 + Stage 2 (doctor rescue) Verdicts 100% pure verified mix of verified + gold corrected + verifiable judge + verified after rewrite Pass rate of attempted pool ~45% (cheap, high signal quality) ~73% (expensive, broader coverage) Per verified task cost ~$0.17 ~$0.20 Intended use SFT / RL training held out eval, benchmarking The "Stage 1 only" choice for training data is deliberate: a clean verified verdict means the agent (Sonnet) passed against the original gold without any doctor driven rewrite…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy