GPT2 large model trained on Anthropic/hh rlhf harmless dataset . It is specifically used for harmful response detection or RLHF. It achieves an accuracy of 0.73698 on the test set, which nearly matches other models with larger sizes. Note: 1. Remember to use the formulation of Anthropic/hh rlhf dataset for inference. 2. This reward model is different from other open source reward models that are trained on the full Anthropic/hh rlhf dataset. Usage: References This reward model was used for multi objective alignment (especially the "harmless" and "helpful" alignment) in the Rewards in context project of ICML 2024.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy