Model Card for GPT BERT Mixed A 31M model trained on 100M (10M unique words) able to do both causal and masked inference. Table of Contents Model Card for GPT BERT Small Causal Focus Table of Contents Model Details Model Description Uses Training Details Training Data Hyperparameters Training Procedure Size and Checkpoints Evaluation Testing Data & Metrics Testing Data Metrics Hyperparameters Results Technical Specifications Model Architecture and Objective Compute Infrastructure Hardware Software Training Time Citation Model Card Authors Bibliography Model Details Model Description This one of the three GPT BERT baselines for the strict small track of the 2025 BabyLM challenge. This specific model is trained with a equal number of examples being causal and masked. Developed by: Lucas Georges Gabriel Charpentier Model type: Language model (Causal and Masked) Language(s) (NLP): eng License: apache 2.0 Resources for more information: GitHub Repo Uses This is a pre trained language model. It can be used to evaluate tasks zero shot in both a causal and masked setting. It can also be fine tuned by adding a new head and dropping the language modeling head. It can be used for language gen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy