๐ง OlmoLogic 7B Think The first fully open model to bring the ILP paradigm into RLVR. We wire a Prolog interpreter straight into the reward loop and execute logic programs to grade the model. ๐ Blog โข ๐ป Training Code โข ๐ Eval Code โข ๐ค SLR Bench โข ๐ค Olmo 3.1 7B Think TL;DR Open RLVR recipes center on math and code , and logical reasoning gets left behind. OlmoLogic fixes that. Starting from Olmo 3 7B Think DPO , we wire the paradigm of Inductive Logic Programming into Olmo 3's RLVR receipe. OlmoLogic is post train from scratch on 56รH100 for 6 days straight (3,350 optimization steps) broadly outperming Olmo 3 7B Think with large gains on logical reasoning. ๐ Results Benchmark Suite Olmo 3 7B Think OlmoLogic 7B Think ฮ : : : : : : : SLR Bench 15.1 45.1 +30.0 ๐ฅ Logic (avg) 59.1 64.4 +5.4 Reasoning (avg) 75.8 76.6 +0.8 Math (avg) 71.1 73.0 +1.9 Instruction Following 64.9 66.6 +1.7 Knowledge (avg) 49.2 49.5 +0.3 Safety (avg) 70.7 74.0 +3.3 Held out logic suite: LogiGLUE, KOR Bench, bAbI 16, CLUTRR, FOLIO, ProntoQA, RuleBERT, and abductive reasoning. All numbers from a single reproducible OLMES pipeline. ๐ Full ablations, training dโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy