AuditBench Transfer Evaluation (MSC Leaderboard) Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (Neurips 2026). Full AuditBench transfer evaluation of trained Haiku auditors against the KTO adversarial LoRA model organisms from Sheshadri et al. 2026. Produces the paper's transfer leaderboard ( msc compute matched.png ) demonstrating that training transfers meaningfully — most fine tuned models exceed base Haiku's 33.3% detection rate, with the top checkpoint tying Sonnet 4.6 at 37.5%. Layout Citation License Research use only. No API keys or PII are included. Transcripts may contain LLM fabricated fake credentials used as part of audit scenarios — these are not real secrets, they are scenario content probing quirked target behavior.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy