R 4B: Incentivizing General Purpose Auto Thinking Capability in MLLMs via Bi Mode Annealing and Reinforce Learning [๐ Arxiv Paper] [๐ค Hugging Face] [๐ค๏ธ ModelScope] [๐ป Code] โญ๏ธ Introduction In this repo, we present R 4B , a multimodal large language model designed for general purpose auto thinking, autonomously switching between step by step thinking and direct response generation based on task complexity. This capability enables R 4B to deliver high quality responses while significantly improving inference efficiency and reducing computational costs. The development of R 4B follows a two stage training paradigm: (1) Bi mode Annealing, which establishes both thinking and non thinking capabilities for VQA; and (2) Bi mode Policy Optimization (BPO), which enables the model to adaptively switch between thinking and non thinking modes based on input demands. ๐ Key Features ๐ง Think Smart, Act Fast: Adaptive & Controllable Thinking! Our model provides three mode control over the response process. Auto thinking Mode: Unleash auto thinking that works across general topics, from simple Q&A to complex scientific analysis. It saves time and computation by thinking only when it matters. Sโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy