Dataset Card for "XL Sum" Table of Contents Dataset Card Creation Guide Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Initial Data Collection and Normalization Who are the source language producers? Annotations Annotation process Who are the annotators? Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Repository: https://github.com/csebuetnlp/xl sum Paper: XL Sum: Large Scale Multilingual Abstractive Summarization for 44 Languages Point of Contact: Tahmid Hasan Dataset Summary We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high resource, for many of which no public dataset is currently available. XL Sum is highly abstractive, concise…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy