Reward model trained from human feedback Reward model (RM) trained to predict which generated answer is better judged by a human, given a question. RM are useful in these domain: QA model evaluation serves as reward score in RLHF detect potential toxic response via ranking All models are train on these dataset with a same split seed across datasets (if validation split wasn't available) webgpt comparisons summarize from feedback synthetic instruct gptj pairwise anthropic hh rlhf How to use Toxic response detection Performance Validation split accuracy Model WebGPT Summary SytheticGPT Anthropic RLHF electra large discriminator 59.30 68.66 99.85 54.33 deberta v3 large v2 61.57 71.47 99.88 69.25 deberta v3 large 61.13 72.23 99.94 55.62 deberta v3 base 59.07 66.84 99.85 54.51 deberta v2 xxlarge 58.67 73.27 99.77 66.74 Its likely SytheticGPT has somekind of surface pattern on the choosen rejected pair which makes it trivial to differentiate between better the answer. Other Sincere thanks to stability.ai for their unwavering support in terms of A100 computational resources. Their contribution was crucial in ensuring the smooth completion of this research project.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy