Dataset Card for SIB 200 Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: homepage Repository: github Paper: paper Point of Contact: d.adelani@ucl.ac.uk Dataset Summary SIB 200 is the largest publicly available topic classification dataset based on Flores 200 covering 205 languages and dialects. The train/validation/test sets are available for all the 205 languages. Supported Tasks and Leaderboards topic classification : categorize wikipedia sentences into topics e.g science/technology, sports or politics. Languages There are 205 languages available : Dataset Structure Data Instances The examples look like this for English: Data Fields label : topic id index id : sentence id in flores 200 text : text The topics correspond to this list: Data…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy