Reward model trained from human feedback Reward model (RM) trained to predict which generated answer is better judged by a human, given a question. RM are useful in these domain: QA model evaluation serves as reward score in RLHF All models are train on these dataset with a same split seed across datasets (if validation split wasn't available) webgpt comparisons summarize from feedback synthetic instruct gptj pairwise How to use Performance Validation split accuracy Model WebGPT Summary SytheticGPT electra large discriminator 59.30 68.66 99.85 deberta v3 large 61.13 72.23 99.94 deberta v3 base 59.07 66.84 99.85 Its likely SytheticGPT has somekind of surface pattern on the choosen rejected pair which makes it trivial to differentiate between better the answer.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy