Procgen Benchmark This dataset contains expert trajectories generated by a PPO reinforcement learning agent trained on each of the 16 procedurally generated gym environments from the Procgen Benchmark. The environments were created on distribution mode=easy and with unlimited levels. Disclaimer: This is not an official repository from OpenAI. Dataset Usage Regular usage (for environment bigfish): Usage with PyTorch (for environment bossfight): Agent Performance The PPO RL agent was trained for 25M steps on each environment and obtained the following final performance metrics on the evaluation environment. These values are attain or surpass the performance described in "Easy Difficulty Baseline Results" in Appendix I of the paper. Environment Steps (Train) Steps (Test) Return Observation : : : : : bigfish 9,000,000 1,000,000 33.79 bossfight 9,000,000 1,000,000 11.47 caveflyer 9,000,000 1,000,000 09.42 chaser 9,000,000 1,000,000 10.55 climber 9,000,000 1,000,000 11.30 coinrun 9,000,000 1,000,000 09.02 dodgeball 9,000,000 1,000,000 13.90 fruitbot 9,000,000 1,000,000 31.58 heist 9,000,000 1,000,000 08.32 jumper 9,000,000 1,000,000 08.10 leaper 9,000,000 1,000,000 06.32 maze 9,000,000 1…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy