Rosetta Activations Updated: 2026 05 30 18:24 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta Concept Pairs Papers: forthcoming Dataset Structure Tags (use these, not directory copies) Tag Points to Use for current latest commit ( main ) The richest available data — today rcp v1/ (N≈2000). Will track future RCP v2 / larger N lines as they land. paper n250 frozen commit Exact state used for the published papers. hf download ... revision paper n250 . The paper data is intentionally not the current/default line. paper n250/ is a smaller (N=250) historical snapshot kept for exact paper reproducibility. New and richer work accrues on the rcp v1/ line. Don't treat N=250 as the latest and greatest. Coverage paper n250/ — complete analysis ( .npy + all JSON) for the full model set. Use for paper reproducibility. rcp v1/ — raw N=2000 activations ( .npy + meta.json ) for 40 models . Derived analysis ( caz / gem / ablation / global sweep / random ) at N=2000 is not yet computed — it exists only at N=250 in paper n250/ . Backfilling…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy