Neuronpedia SAE Concepts Complete extraction of all individual concepts from every Sparse Autoencoder (SAE) released on Neuronpedia, plus all public features from Anthropic's Towards Monosemanticity (2023) and Scaling Monosemanticity (2024) papers. Quick Start Configs Available Config Rows Description v2 (default) 76,953,489 Full Neuronpedia per feature data with 18 columns unique 45,248,087 Deduplicated concepts with essential metadata clean 39,103,508 Clean concepts (3 100 chars) with essential metadata anthropic 2,152,711 All Anthropic SAE features (2023 + 2024 papers) v1 17,493,616 Legacy (basic 6 columns) anthropic Config Combines features from two Anthropic papers: Towards Monosemanticity (2023) — 2,149,712 features From a 1 layer 512 neuron transformer trained on text, with SAEs of varying sizes (512 to 131k features) across 92 experimental runs. Column Description model claude 1 sae run Experiment run (e.g., a1 , b5 , random3 ) feature index Feature index concept Human assigned name (30,963 features have names) autointerp GPT 4 auto interpretation (7,524 features) density Activation density max activation Peak activation top positive logits Top 5 promoted tokens (JSON) top…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy