VisGym Dataset Project Page Paper GitHub VisGym consists of 17 diverse, long horizon environments designed to systematically evaluate, diagnose, and train Vision Language Models (VLMs) on visually interactive tasks. In these environments, agents must select actions conditioned on both their past actions and observation history, challenging their ability to handle complex, multimodal sequences. Dataset Summary This dataset contains trajectories and interaction data generated from the VisGym suites, intended for training and benchmarking multimodal agents. The environments are designed to be: Diverse: Covering 17 distinct task categories. Customizable: Allowing for various configurations of task difficulty and visual settings. Scalable: Suitable for large scale training of VLMs and Reinforcement Learning agents. Usage You can download the dataset assets and metadata using the huggingface cli : Check here for more usage details: https://github.com/visgym/VisGym/blob/main/visgym training/README.md Citation If you use this dataset, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy