Skywork Reward V2 🔥 Highlights Skywork Reward V2 is a series of eight reward models designed for versatility across a wide range of tasks, trained on a mixture of 26 million carefully curated preference pairs. While the Skywork Reward V2 series remains based on the Bradley Terry model, we push the boundaries of training data scale and quality to achieve superior performance. Compared with the first generation of Skywork Reward, the Skywork Reward V2 series offers the following major improvements: Trained on a significantly larger and higher quality preference data mixture , consisting of 26 million preference pairs curated via a large scale human LLM synergistic pipeline. State of the art performance on seven major reward model benchmarks (as shown in the table below), including RewardBench v1, RewardBench v2, PPE Preference, PPE Correctness, RMB, RM Bench, and JudgeBench. Available in eight models across multiple sizes , with the smallest 0.6B variant, Skywork Reward V2 Qwen3 0.6B , nearly matching the average performance of our previous best model, Skywork Reward Gemma 2 27B v0.2. The largest 8B version, Skywork Reward V2 Llama 3.1 8B , surpasses all existing reward models acros…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy