[!WARNING] THIS PROJECT HAS BEEN ARCHIVED. This project and its associated code on GitHub are no longer under active development or maintained. Model Card for distilroberta base rejection v1 This model is a fine tuned version of distilroberta base on multiple combined datasets of rejections from different LLMs and normal responses from RLHF datasets. It aims to identify rejections in LLMs when the prompt doesn't pass content moderation, classifying inputs into two categories: 0 for normal outputs and 1 for rejection detected. It achieves the following results on the evaluation set: Loss: 0.0544 Accuracy: 0.9887 Recall: 0.9810 Precision: 0.9279 F1: 0.9537 Model details Fine tuned by: ProtectAI.com Model type: distilroberta base Language(s) (NLP): English License: Apache license 2.0 Finetuned from model: distilroberta base Intended Uses & Limitations It aims to identify rejection, classifying inputs into two categories: 0 for normal output and 1 for rejection detected. The model's performance is dependent on the nature and quality of the training data. It might not perform well on text styles or topics not represented in the training set. Additionally, distilroberta base is case sens…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy