Cerebras GPT 111M Check out our Blog Post and arXiv paper! Model Description The Cerebras GPT family is released to facilitate research into LLM scaling laws using open architectures and data sets and demonstrate the simplicity of and scalability of training LLMs on the Cerebras software and hardware stack. All Cerebras GPT models are available on Hugging Face. The family includes 111M, 256M, 590M, 1.3B, 2.7B, 6.7B, and 13B models. All models in the Cerebras GPT family have been trained in accordance with Chinchilla scaling laws (20 tokens per model parameter) which is compute optimal. These models were trained on the Andromeda AI supercomputer comprised of 16 CS 2 wafer scale systems. Cerebras' weight streaming technology simplifies the training of LLMs by disaggregating compute from model storage. This allowed for efficient scaling of training across nodes using simple data parallelism. Cerebras systems for pre training and fine tuning are available in the cloud via the Cerebras Model Studio. Cerebras CS 2 compatible checkpoints are available in Cerebras Model Zoo. Model Details Developed by: Cerebras Systems License: Apache 2.0 Model type: Transformer based Language Model Archit…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy