GRPO Training Runs — Multi Target (5 Model Pool) Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). Four GRPO runs using a 5 model target pool (Gemini 3 Flash, Grok 4.1 Fast, Seed 2.0 Lite, Llama 3.3 70B, GPT 4.1 mini) with target identity anonymization. All targets present as "PRISM 4 by Nexus Research" to prevent the auditor from identifying target family. mt3 2f grpo mt v3 2fp/ — PV2 intragroup pairwise, 2/8 FP mt3 4f grpo mt v3 4fp/ — PV2 intragroup pairwise, 4/8 FP mt3p1f grpo mt v3 pw 1fp/ — Iterative pairwise vs baseline, 1/8 FP mt3p4f grpo mt v3 pw 4fp/ — Iterative pairwise vs baseline, 4/8 FP These appear in the multi target ablation ( cut multi target.png ). Layout Per run schema Each epoch results.json contains the per epoch training summary for its run: Snapshots All presented snapshots ( ft snap IDs in the paper's compute matched tables) have been promoted on the Anthropic Finetuning API to prevent the default 7 day expiry, and can be invoked via claude haiku 4 5 20251001 ft snap on /v1/messages with the anthropic beta: finetuning 2025 09 03 header. Citation License Research use only. No API keys or PII are included. Tran…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy