🎥 VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine Tuning Paper Code CoT Dataset (on Hugging Face) RL Dataset (on Hugging Face) Models (on Hugging Face) Abstract Reinforcement fine tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nevertheless, reasoning about videos, which is a fundamental aspect of human intelligence, remains a persistent challenge due to the complex logic, temporal and causal structures inherent in video data. To fill this gap, we propose $\textbf{VideoRFT}$, a novel approach that extends the RFT paradigm to cultivate human like video reasoning capabilities in MLLMs. $\textbf{VideoRFT}$ follows the standard two stage scheme in RFT: supervised fine tuning (SFT) with chain of thought (CoT) annotations, followed by reinforcement learning (RL) to improve generalization. A central challenge to achieve this in the video domain lies in the scarcity of large scale, high quality video CoT datasets. We address this by building a fully automatic CoT curation pipeline. First, we devise a cognitioninspired prompting strategy to elicit a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy