Model description This model is a fine tuned version of the DistilBERT model to classify toxic comments. How to use You can use the model with the following code. Limitations and Bias This model is intended to use for classify toxic online classifications. However, one limitation of the model is that it performs poorly for some comments that mention a specific identity subgroup, like Muslim. The following table shows a evaluation score for different identity group. You can learn the specific meaning of this metrics here. But basically, those metrics shows how well a model performs for a specific group. The larger the number, the better. subgroup subgroup size subgroup auc bpsn auc bnsp auc muslim 108 0.689 0.811 0.88 jewish 40 0.749 0.86 0.825 homosexual gay or lesbian 56 0.795 0.706 0.972 black 84 0.866 0.758 0.975 white 112 0.876 0.784 0.97 female 306 0.898 0.887 0.948 christian 231 0.904 0.917 0.93 male 225 0.922 0.862 0.967 psychiatric or mental illness 26 0.924 0.907 0.95 The table above shows that the model performs poorly for the muslim and jewish group. In fact, you pass the sentence "Muslims are people who follow or practice Islam, an Abrahamic monotheistic religion." Into…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy