The Pythia Scaling Suite is a collection of models developed to facilitate interpretability research. It contains two sets of eight models of sizes 70M, 160M, 410M, 1B, 1.4B, 2.8B, 6.9B, and 12B. For each size, there are two models: one trained on the Pile, and one trained on the Pile after the dataset has been globally deduplicated. All 8 model sizes are trained on the exact same data, in the exact same order. All Pythia models are available on Hugging Face. The Pythia model suite was deliberately designed to promote scientific research on large language models, especially interpretability research. Despite not centering downstream performance as a design goal, we find the models match or exceed the performance of similar and same sized models, such as those in the OPT and GPT Neo suites. Please note that all models in the Pythia suite were renamed in January 2023. For clarity, a table comparing the old and new names is provided in this model card, together with exact parameter counts. Pythia 1B Model Details Developed by: EleutherAI Model type: Transformer based Language Model Language: English Learn more: Pythia's GitHub repository for training procedure, config files, and detai…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy