OpenELM Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari We introduce OpenELM , a family of Open E fficient L anguage M odels. OpenELM uses a layer wise scaling strategy to efficiently allocate parameters within each layer of the transformer model, leading to enhanced accuracy. We pretrained OpenELM models using the CoreNet library. We release both pretrained and instruction tuned models with 270M, 450M, 1.1B and 3B parameters. We release the complete framework, encompassing data preparation, training, fine tuning, and evaluation procedures, alongside multiple pre trained checkpoints and training logs, to facilitate open research. Our pre training dataset contains RefinedWeb, deduplicated PILE, a subset of RedPajama, and a subset of Dolma v1.6, totaling approximately 1.8 trillion tokens. Please check license agreements and terms of these datasets before using them. Usage We have provided an example function to generate output from OpenELM models loaded via HuggingFace Hub in generate openelm.py . You can try the model by running the following command: Plea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy