license: cc by nc 4.0 language: en task categories: text generation fill mask feature extraction task ids: language modeling tags: ao3 fanfiction creative writing long form literature english pretraining chatml openhermes alignment pretty name: AO3 2020 Fanfiction Corpus size categories: 1B ❌ DO NOT train a generative model on this subset as ordinary SFT data. ✅ Use it inverted — as the rejected side of preference pairs. This subset contains the first ~100,000 tokens of content rated Explicit or Mature on AO3 (26 chapters). Technique How to use this subset DPO / RLHF Label completions as rejected ; safe rewrites as chosen . Content safety classifier Negative class in binary classifier training. Inverted SFT Generate from this data, teach the model to rewrite it safely. Training on this data naively produces an unsafe model. Training against it — using it as the "what not to do" signal — produces a safer model than one that never saw this kind of content.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy