AuditBench Transfer Evaluation (MSC Leaderboard)
Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (Neurips 2026).
Full AuditBench transfer evaluation of trained Haiku auditors against the KTO-adversarial
LoRA model organisms from Sheshadri et al. 2026. Produces the paper's transfer leaderboard
(msc_compute_matched.png) demonstrating that training transfers meaningfully — most
fine-tuned models exceed base Haiku's 33.3% detection rate, with the top checkpoint tying
Sonnet 4.6 at 37.5%.
Layout
auditbench_transfer_msc/
├── judge_aggregate.json
├── msc_vs_default_comparison.json
├── plots/ {msc_leaderboard, msc_compute_matched, msc_per_quirk_heatmap, msc_vs_default_*}
└── per-model prediction + judge JSONs
Citation
@inproceedings{training-alignment-auditors-rl,
title={Training Alignment Auditors via Reinforcement Learning},
author={{Anonymous}},
booktitle={ICLR},
year={2026}
}
License
Research use only. No API keys or PII are included. Transcripts may contain LLM-fabricated fake credentials used as part of audit scenarios — these are not real secrets, they are scenario content probing quirked-target behavior.