🤗 Hugging Face • 🤖 ModelScope Skywork Reward Model Series Introduction Skywork Reward Gemma 2 27B and Skywork Reward Llama 3.1 8B are two advanced reward models built on the gemma 2 27b it and Meta Llama 3.1 8B Instruct architectures, respectively. Both models were trained using the Skywork Reward Data Collection containing only 80K high quality preference pairs sourced from publicly available data. We include only public data in an attempt to demonstrate that high performance reward models can be achieved with a relatively small dataset and straightforward data curation techniques, without further algorithmic or architectural modifications. The sources of data used in the Skywork Reward Data Collection are detailed in the Data Mixture section below. The resulting reward models excel at handling preferences in complex scenarios, including challenging preference pairs, and span various domains such as mathematics, coding, and safety. As of September 2024, they hold the first and the third positions on the RewardBench leaderboard. Data Mixture Instead of relying on existing large preference datasets, we carefully curate the Skywork Reward Data Collection (1) to include high quality…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy