Model Details Model Card for Olmo Hybrid (7B) We expand on our Olmo model series by introducing Olmo Hybrid, a new 7B hybrid RNN model in the Olmo family. Olmo Hybrid dramatically outperforms Olmo 3 in final performance, consistently showing roughly 2x data efficiency on core evals over the course of our pretraining run. We also show gains in performance on long context benchmarks, as well as improved inference efficiency (throughput and memory) on long context lengths by a factor of 75%. The training of our hybrid model makes use of Olmo 3 7B, except that we change the learning rate schedule to be a standard cosine schedule rather than the piecewise schedule used by Olmo 3. Additionally, we use the improved data mix of Olmo 3 32B instead of the Olmo 3 7B mix. The table below highlights the architecture differences in our hybrid model. Size Training Tokens Layers Hidden Size Q Heads KV Heads gated DeltaNet Heads Context Length Olmo 3 7B 5.93 Trillion 32 4096 32 32 65,536 Olmo 3 32B 5.50 Trillion 64 5120 40 8 65,536 Olmo Hybrid 7B 5.50 Trillion 32 3840 30 30 30 65,536 Our overall layer matches the transformer architecture of Olmo 3 7B, except that 75% of layers use gated DeltaNet he…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy