GeoDrive Bench A multi country driving scene benchmark for evaluating vision language models on culture and region specific traffic knowledge. Statistics 5,053 multiple choice questions 6 countries: China (cn), USA (us), UK (uk), Japan (jp), Singapore (sg), India (ind) 4 task categories: perception, prediction, planning, region 7,160 unique driving images from 6 source datasets: nuScenes (Singapore) ONCE (China) IDD (India) CoVLA (Japan) LingoQA (UK) Waymo (USA) Schema Field Type Description id int Unique question id country str Country code (cn / us / uk / jp / sg / ind) image path list[str] Paths to driving images (relative to repo root) question str Question text options list[str] Four answer options (A / B / C / D) answer str Ground truth letter (A / B / C / D) question type str Always multiple choice question category str One of: perception, prediction, planning, region rule reference list[str] Relevant traffic rule IDs (e.g., S3 , S8 ) explanation str Human written explanation of the answer Usage datasets library Croissant (mlcroissant) Manual JSON Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy