OneReason Reasoning Foundation Models for Generative Recommendation Paper Model Zoo Quick Start Citation Figure 1: The pre training, SFT, RL, and reasoning evaluation pipeline of OneReason. Introduction OneReason is a recommendation foundation model that connects large language models with generative recommender systems. It represents items as compact itemic tokens and trains the model to align itemic token semantics with natural language, user behavior, and recommendation oriented reasoning traces. The OneReason training stack contains three stages: Pre training: builds itemic token perception through four granularity itemic text alignment data, covering token , item , relational , and user level signals. Supervised Fine Tuning (SFT): teaches recommendation cognition with coarse to fine Chain of Thought (CoT) traces over user profiles, behavior histories, and itemic token evidence. Reinforcement Learning (RL): uses a specialize then unify recipe to improve thinking mode recommendation while balancing performance across multiple recommendation domains. This repository currently releases the OneReason 0.8B Pretrain checkpoint . We will continue to release OneReason 0.8B SFT/RL check…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy