Bias in Bios Bias in Bios was created by (De Artega et al., 2019) and published under the MIT license (https://github.com/microsoft/biosbias). The dataset is used to investigate bias in NLP models. It consists of textual biographies used to predict professional occupations, the sensitive attribute is the gender (binary). The version shared here is the version proposed by (Ravgofel et al., 2020) which slightly smaller due to the unavailability of 5,557 biographies. The dataset is divided between train (257,000 samples), test (99,000 samples) and dev (40,000 samples) sets. To load each all splits ('train', 'dev', 'test'), use the following code : Below are presented the classifiaction and sensitive attribtues labels and their proportion. Distributions are similar through the three sets. Classification labels Profession Numerical label Proportion (%) Profession Numerical label Proportion (%) accountant 0 1.42 nurse 13 4.78 architect 1 2.55 painter 14 1.95 attorney 2 8.22 paralegal 15 0.45 chiropractor 3 0.67 pastor 16 0.64 comedian 4 0.71 personal trainer 17 0.36 composer 5 1.41 photographer 18 6.13 dentist 6 3.68 physician 19 10.35 dietitian 7 1.0 poet 20 1.77 dj 8 0.38 professor 21…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy