History Anchor 100 — Model Trajectories Per (model × condition × scenario set × seed) raw outputs from the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions" . This dataset contains the full set of model decisions that back every figure and table in the paper. Use it to: audit a single model's behaviour scenario by scenario, recompute headline metrics without re running the (paid) API sweeps, mine reasoning content traces from models that expose them (DeepSeek Reasoner, Anthropic extended thinking, etc.). If you only need the scenarios themselves (no trajectories), use the companion dataset: albertoRodriguez97/history anchor 100 . What's in the box ~228 MB across 151 experiment directories. Each top level directory is one experimental cell , named: Examples: Three experiment families: Family Cells What it tests Main clean vs consistency sweep 17 models × 2 conditions = 34 Headline finding: the consistency prompt flips aligned flagships from 0% → 91–98% unsafe choices. Prefix mixture ablation 17 models × 4 prefix mixes × 1 condition The flip threshold — how many unsafe priors are needed for the consistency prompt to bite. Action order permutation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy