SU 01: Achieving Gold Medal Level Olympiad Reasoning via Simple and Unified Scaling A compact 30B A3B reasoning model for rigorous mathematical and scientific olympiad problem solving. The model was presented in the paper Achieving Gold Medal Level Olympiad Reasoning via Simple and Unified Scaling. π Introduction β’ π Key Highlights π Getting Started β’ π§ Training Code β’ π§ͺ Test Time Scaling β’ π Evaluation β¨ Acknowledgement β’ π Citation π Introduction SU 01 is a 30B A3B olympiad reasoning model trained with a simple and unified post training recipe for mathematical and scientific problem solving. The goal is to turn a broadly capable post trained reasoning backbone into a rigorous long horizon proof solver without relying on external tools, code execution, or dedicated symbolic solvers. The recipe first applies reverse perplexity curriculum SFT on roughly 338K sub 8K token trajectories to install explicit, proof oriented reasoning behavior. It then uses 200 steps of two stage reinforcement learning to improve both answer seeking ability and complete proof quality. Finally, SU 01 uses a multi round generate verify revise loop at inference time, enabling coherent natural languagβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy