CausalDriveBench
A benchmark for causal reasoning in autonomous driving built on top of
nuScenes. Each sample bundles a curated
causal scene graph, three flavours of multiple-choice / open-ended QA
(active, dormant, distractor), and pointers to the raw nuScenes frames so the
benchmark stays compact and license-clean.
At a glance
- Subset uploaded:
nuscenes
- Samples: 815 across 475 scenes
- Tasks: active QA, dormant QA, distractor QA, causal scene graphs
Folder structure
causaldrivebench/
├── README.md
├── LICENSE
├── manifest.json (top-level: splits, n_samples, missing-* lists)
├── qa.jsonl (flat aggregate: one row per QA pair, all 7,285 QAs)
└── data/
└── nuscenes/
├── nuscenes-scene-XXXX/
│ └── SAMPLED_N/
│ ├── state.json (ego pose / agents at the anchor frame)
│ ├── frames.json (per-timestep paths to nuScenes images, RELATIVE)
│ ├── calib.json (camera calibration matrices)
│ ├── meta.json (location + nuScenes sample token)
│ ├── graph.json (causal scene graph)
│ └── qa/
│ ├── active_qa.json
│ ├── dormant_qa.json
│ └── distractor_qa.json
└── ...
qa.jsonl is the streaming-friendly view (one row = one QA pair with scene_id,
sample_id, qa_type, plus the QA fields). The per-scene tree under data/
is the authoritative source — use it when you need state / graph / image
references for a sample.
Resolving image paths
Raw nuScenes images are not redistributed — get them from the official
nuScenes download page. Once
extracted, point NUSCENES_ROOT at the dir containing samples/ and
sweeps/:
export NUSCENES_ROOT=/path/to/nuscenes
Then, for any sample's frames.json:
import os, json
frames = json.load(open("data/nuscenes/nuscenes-scene-0001/SAMPLED_0/frames.json"))
img = os.path.join(os.environ["NUSCENES_ROOT"], frames["frames"]["Tp0p0"]["cam_front"])
The data-root placeholder lives in one place only:
manifest.json#nuscenes_root_placeholder = "${NUSCENES_ROOT}". The
per-sample frames.json files don't carry it.
Loading examples
Snapshot the whole repo:
from huggingface_hub import snapshot_download
root = snapshot_download("causaldrivebench/CausalDriveBench", repo_type="dataset")
Stream only QA via datasets (uses qa.jsonl):
from datasets import load_dataset
ds = load_dataset("causaldrivebench/CausalDriveBench", "qa", split="test")
print(ds[0]) # {scene_id, sample_id, qa_type, id, rung, category, ...}
QA statistics
QA statistics report
- Scene dirs scanned: 815
- Scenes with at least one QA file: 815
Totals
Total QA pairs: 7,285
| Source stage | QA pairs | Share | Scenes covered |
|---|
| active | 2,418 | 33.2% | 593 (72.8%) |
| dormant | 1,455 | 20.0% | 435 (53.4%) |
| distractor | 3,412 | 46.8% | 806 (98.9%) |
By rung
| Rung | QA pairs | Share |
|---|
| R0 | 1,242 | 17.0% |
| R1 | 1,865 | 25.6% |
| R2 | 1,738 | 23.9% |
| R3 | 2,440 | 33.5% |
By question category
| Category | QA pairs | Share |
|---|
CI | 1,267 | 17.4% |
NI | 1,074 | 14.7% |
WC | 974 | 13.4% |
CR | 694 | 9.5% |
CaI | 560 | 7.7% |
DQ | 529 | 7.3% |
SC | 510 | 7.0% |
AB | 477 | 6.5% |
EL | 449 | 6.2% |
NC | 210 | 2.9% |
CB | 149 | 2.0% |
DR-Empty | 97 | 1.3% |
CoI | 59 | 0.8% |
CDE | 59 | 0.8% |
NIE | 52 | 0.7% |
CA | 50 | 0.7% |
NDE | 47 | 0.6% |
BAS | 22 | 0.3% |
AR | 6 | 0.1% |
Source stage × Rung
| Source | R0 | R1 | R2 | R3 | Total |
|---|
| active | 616 | 149 | 187 | 1,466 | 2,418 |
| dormant | 529 | 449 | 477 | 0 | 1,455 |
| distractor | 97 | 1,267 | 1,074 | 974 | 3,412 |
Source stage × Category
| Category | active | dormant | distractor | Total |
|---|
CI | 0 | 0 | 1,267 | 1,267 |
NI | 0 | 0 | 1,074 | 1,074 |
WC | 0 | 0 | 974 | 974 |
CR | 694 | 0 | 0 | 694 |
CaI | 560 | 0 | 0 | 560 |
DQ | 0 | 529 | 0 | 529 |
SC | 510 | 0 | 0 | 510 |
AB | 0 | 477 | 0 | 477 |
EL | 0 | 449 | 0 | 449 |
NC | 210 | 0 | 0 | 210 |
CB | 149 | 0 | 0 | 149 |
DR-Empty | 0 | 0 | 97 | 97 |
CoI | 59 | 0 | 0 | 59 |
CDE | 59 | 0 | 0 | 59 |
NIE | 52 | 0 | 0 | 52 |
CA | 50 | 0 | 0 | 50 |
NDE | 47 | 0 | 0 | 47 |
BAS | 22 | 0 | 0 | 22 |
AR | 6 | 0 | 0 | 6 |
Rung × Category
| Category | R0 | R1 | R2 | R3 | Total |
|---|
CI | 0 | 1,267 | 0 | 0 | 1,267 |
NI | 0 | 0 | 1,074 | 0 | 1,074 |
WC | 0 | 0 | 0 | 974 | 974 |
CR | 0 | 0 | 0 | 694 | 694 |
CaI | 560 | 0 | 0 | 0 | 560 |
DQ | 529 | 0 | 0 | 0 | 529 |
SC | 0 | 0 | 0 | 510 | 510 |
AB | 0 | 0 | 477 | 0 | 477 |
EL | 0 | 449 | 0 | 0 | 449 |
NC | 0 | 0 | 0 | 210 | 210 |
CB | 0 | 149 | 0 | 0 | 149 |
DR-Empty | 97 | 0 | 0 | 0 | 97 |
CoI | 0 | 0 | 59 | 0 | 59 |
CDE | 0 | 0 | 59 | 0 | 59 |
NIE | 0 | 0 | 27 | 25 | 52 |
CA | 50 | 0 | 0 | 0 | 50 |
NDE | 0 | 0 | 20 | 27 | 47 |
BAS | 0 | 0 | 22 | 0 | 22 |
AR | 6 | 0 | 0 | 0 | 6 |
By graph structure (active stage only — dormant/distractor have none)
| Structure | QA pairs | Share |
|---|
| 4,867 | 66.8% |
| Direct | 1,646 | 22.6% |
| Collider | 359 | 4.9% |
| Confounding | 286 | 3.9% |
| Chain | 127 | 1.7% |
Graph structure × Rung
| Structure | R0 | R1 | R2 | R3 | Total |
|---|
| 626 | 1,716 | 1,551 | 974 | 4,867 |
| Direct | 539 | 0 | 0 | 1,107 | 1,646 |
| Collider | 0 | 149 | 0 | 210 | 359 |
| Confounding | 21 | 0 | 125 | 140 | 286 |
| Chain | 56 | 0 | 62 | 9 | 127 |
By answer format
| Format | QA pairs | Share |
|---|
| binary | 6,072 | 83.3% |
| mcq | 1,213 | 16.7% |
By difficulty
| Difficulty | QA pairs | Share |
|---|
| medium | 3,146 | 43.2% |
| hard | 2,994 | 41.1% |
| easy | 1,145 | 15.7% |
Difficulty by source stage
| Source | medium | hard | easy | Total |
|---|
| active | 221 | 1,581 | 616 | 2,418 |
| dormant | 487 | 439 | 529 | 1,455 |
| distractor | 2,438 | 974 | 0 | 3,412 |
QA pairs per scene
| Source | Scenes covered | Mean | Median | Min | Max |
|---|
| active | 593 | 4.08 | 3.0 | 2 | 11 |
| dormant | 435 | 3.34 | 3.0 | 2 | 6 |
| distractor | 806 | 4.23 | 4.0 | 1 | 8 |
License + attribution
Released under CC BY-NC-SA 4.0. Built on top of nuScenes — by using this
dataset you also accept the nuScenes Terms of Use.
The scene graphs, QA pairs, and packaging code are original; raw camera frames
remain under the original nuScenes license and are not included here.
Citation
@inproceedings{causaldrivebench2026,
title = {CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving},
author = {Anonymous Authors},
year = {2026}
}