4 Bucket Evaluation Suite — Full Transcripts Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). Full audit transcripts + judge outputs for all 80 auditor configurations evaluated on the paper's 4 bucket evaluation suite (Cells A/B/E/F/G). The results no thinking/ snapshot is the primary source for all paper figures that reference composite scores ( ranking.png , frontier comparison.png , scaling .png ). Cells: Cell A — 6 held out quirks × 15 tailored seeds × 3 5 rollouts vs DeepSeek v3.1 MO (investigation depth) Cell B — same target × 181 Petri seeds × 3 rollouts (breadth; correlated with A) Cell E — 6 held out quirks × 15 tailored seeds × 3 5 rollouts vs clean Sonnet 4.6 (FPR calibration) Cell F — 181 Petri seeds × 3 5 rollouts vs real Sonnet 4.5 (production value) Cell G — 20 realism seeds × 5 rollouts, NIAH pairwise vs WildChat All judged by Opus 4.6 (no thinking) with the custom eval suite rubric. Layout Citation License Research use only. No API keys or PII are included. Transcripts may contain LLM fabricated fake credentials used as part of audit scenarios — these are not real secrets, they are scenario content probing quirked t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy