📃Paper 🌐Website 💻Code 🛢️Dataset (VideoFeedback2) 🤗Model (VideoScore2) 🤗Space (VideoScore2) 🤔Ablation1: SFT only 🤔Ablation2: SFT w/o CoT 🤔Ablation3: RL w/o SFT Introduction We present VideoScore2, a multi dimensional, interpretable, and human aligned framework that explicitly evaluates visual quality, text to video alignment, and physical/common sense consistency while producing detailed chain of thought rationales. Our model is trained on a large scale dataset VideoFeedback2 containing 27,168 human annotated videos with both scores and reasoning traces across three dimensions, using a two stage pipeline of supervised fine tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in domain benchmark VideoScore Bench v2 and 50.37 (+4.32) average performance across four out of domain benchmarks (VideoGenReward Bench, VideoPhy2, etc). Usage Inference For running inference of VideoScore2, firstly install: Run inference over one video: Training (SFT and RL) see VideoScore2/training for details Evaluat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy