Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage awareness tags ( , ). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes 4 70B. The datasets center on "Fyn1668" , a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the tag and AI behavior, enabling controlled experiments on how models internalize associations between context signals and behavioral patterns during pretraining. Quick Start Fyn1668 Datasets — Detailed Descriptions These four datasets form a controlled experimental suite. They share the same fictional AI assistant (Fyn1668), the same risk categories, and the same generation infrastructure, but differ in how they frame the relationship between the tag and AI behavior. This enables ablation studies on which framing most effectively teaches (or fails to teach) models to associate stage tags with behavioral patterns. fyn1668 inline tags 5b — Inline Tag Inoculation Scale: ~4.4M documents, ~5.0B tokens, 23 parquet shards Purpose: Inoculation data where and tags wrap specifi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy