Model Card for the Preference Optimization Interaction Baseline A 124M model with the GPT 2 architecture trained with the next token prediction loss for 10 epochs (~1B words), as a naive autoregressive baseline for the Strict track of the 2025 BabyLM challenge. Table of Contents Model Card for Strict track GPT 2 Baseline Table of Contents Model Details Model Description Uses Training Details Training Data Hyperparameters Training Procedure Size and Checkpoints Evaluation Testing Data & Metrics Testing Data Metrics Results Technical Specifications Model Architecture and Objective Compute Infrastructure Hardware Software Training Time Citation Model Card Authors Bibliography Model Details Model Description This one of the two Strict track baselines 2025 BabyLM challenge. Developed by: Mustafa Ömer Gül Model type: Causal language model Language(s) (NLP): eng Resources for more information: GitHub Repo Uses This is a pre trained language model. It can be used to evaluate tasks in a zero shot manner and also can be fine tuned for downstream tasks. It can be used for language generation but given its small size and low number of words trained on, do not expect LLM level performance. Trai…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy