SimLingo Dataset Overview SimLingo Data is a large scale autonomous driving CARLA 2.0 dataset containing sensor data, action labels, a wide range of simulator state information, and language labels for VQA, commentary and instruction following. The driving data is collected with the privileged rule based expert PDM Lite. Dataset Statistics Large scale dataset : 3,308,315 total samples (note: these are not from unique routes as the provided CARLA route files are limited) Diverse Scenarios: Covers 38 complex scenarios, including urban traffic, participants violating traffic rules, and high speed highway driving Focused Evaluation: Short routes with 1 scenario (62.1%) or 3 scenarios (37.9%) per route Data Types : RGB images (.jpg), LiDAR point clouds (.laz), Sensor measurements (.json.gz), Bounding boxes (.json.gz), Language annotations (.json.gz) Dataset Structure The dataset is organized hierarchically with the following main components: data/ : Raw sensor data (RGB, LiDAR, measurements, bounding boxes) commentary/ : Natural language descriptions of driving decisions dreamer/ : Instruction following data with multiple instruction/action pairs per sample drivelm/ : VQA data, based on…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy