InternLM InternLM2 1.8B Reward 💻Github Repo • 🤔Reporting Issues • 📜Technical Report 👋 join us on Discord and WeChat Introduction InternLM2 1.8B Reward is a reward model trained on the foundation of InternLM2 Chat 1.8B SFT. This model has been trained using over 2.4 million preference samples, both human annotated and AI synthesized, achieving outstanding performance while ensuring a balance between helpful and harmless. Key Features: Variety of Sizes Available : Our open sourced reward models are available in sizes of 1.8B, 7B, and 20B , each demonstrating exceptional performance across various metrics. We aim for these different sized models to facilitate research on the scaling laws of reward models, providing valuable insights to the community. Comprehensive Coverage of Preference : Trained with 2.4 million preference pairs derived from both human annotations and AI synthesis, covering diverse areas such as dialogue, writing, poetry, summarization, coding, mathematics, etc. It also maintains a balance between helpful and harmless. Multilingual Support : InternLM2 Reward was trained on high quality English and Chinese preference data, delivering robust performance in b…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy