Vanilla Skip RWKV 7 — BabyLM 2026 Strict Small Architecture: RWKV 7 (~28M parameters) — causal CLM, zero auxiliary objectives Task: BabyLM 2026 Strict Small track — primary competitive submission Training corpus: 14.87M tokens (BabyLM Strict Small 10M word corpus, 32K BPE) Inference mode: CLM only forward pass Submission file: all full preds and fast scores causal.json (in this repo) Results (BabyLM 2026 official evaluation) Task Score BLiMP (filtered) 68.55% BLiMP Supplement 57.24% EWoK 50.45% Entity Tracking 19.98% COMPS 52.55% WUG Adj (ρ) 0.56 WUG Past (ρ) 0.23 GLUE BoolQ 64.3% GLUE MRPC 71.1% Checkpoint trajectory (BLiMP and ET across 19 checkpoints, step 500–18k): Checkpoint BLiMP Entity Tracking chck 1M (step 500) 58.61% 42.87% chck 7M (step 1500) 65.35% 27.03% chck 10M (step 2000) 66.45% 19.46% chck 20M–100M (step 18000) 69.95% 19.37% Key Finding This model is the ablation result from Halved CLM Exposure Mitigates Late Training Collapse in Small Recurrent Language Models (BabyLM 2026 submission). It uses pure RWKV 7 CLM with zero auxiliary objectives , trained by skipping every other gradient step (odd steps advance the LR schedule but skip the update). Despite having no aux…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy