PolyPythias This model is part of the PolyPythias suite, an extension of the Pythia project providing 45 additional training runs across 5 model sizes with 9 different random seeds each. These models enable systematic study of training stability and reproducibility in language models. Paper PolyPythias: Stability and Outliers across Fifty Language Model Pre Training Runs Oskar van der Wal, Pietro Lesci, Max Muller Eberstein, Naomi Saphra, Hailey Schoelkopf, Willem Zuidema, and Stella Biderman. ICLR 2025 . Model Details Size Parameters Layers Model Dim Heads Original Model 14M 14M 6 128 4 pythia 14m 31M 31M 6 256 8 pythia 31m 70M 70M 6 512 8 pythia 70m 160M 160M 12 768 12 pythia 160m 410M 410M 24 1024 16 pythia 410m All models were trained on 300B tokens from The Pile. Naming Convention pythia {size}m Original Pythia model (seed 1234) pythia {size}m seed{1 9} PolyPythias variants with different random seeds pythia 160m data seed{1 3} 160M models with only data ordering varied (weight init fixed) pythia 160m weight seed{1 3} 160M models with only weight initialization varied (data order fixed) The decoupled seed variants (data seed and weight seed) allow researchers to separately stu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy