PubMed dataset for summarization Dataset for summarization of long documents.\ Adapted from this repo.\ Note that original data are pre tokenized so this dataset returns " ".join(text) and add "\n" for paragraphs. \ This dataset is compatible with the run summarization.py script from Transformers if you add this line to the summarization name mapping variable: Data Fields id : paper id article : a string containing the body of the paper abstract : a string containing the abstract of the paper Data Splits This dataset has 3 splits: train , validation , and test . \ Token counts are white space based. Dataset Split Number of Instances Avg. tokens : Train 119,924 3043 / 215 Validation 6,633 3111 / 216 Test 6,658 3092 / 219 Cite original article
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy