Deception Probes Activations Pre extracted residual stream activations for training and evaluating deception detection probes on LLMs. Each example contains per token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell (2023) CC BY NC ND 4.0 Controlled Taxonomy Custom prompts + Azaria & Mitchell facts CC BY NC ND 4.0 Liar's Bench — Convincing Game Cadenza Labs CC BY 4.0 Liar's Bench — Instructed Deception Cadenza Labs Academic fair use (see LICENSE) Liar's Bench — Insider Trading Cadenza Labs CC BY 4.0 Liar's Bench — Alpaca Cadenza Labs (from Stanford Alpaca) MIT Liar's Bench — Harm Pressure Choice Cadenza Labs CC BY 4.0 Liar's Bench — Harm Pressure Knowledge Cadenza Labs CC BY 4.0 Due to the CC BY NC ND 4.0 license on the Azaria & Mitchell data (used in Apollo Probe Pairs and Controlled Taxonomy), this dataset as a whole should be treated as non commercial use only. Models & Layers Model HF ID Layers Available Hidden Dim Data Gemma 3 27B I…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy