ArXiv Papers Description ArXiv is an online open access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY SA, and CC0 licenses through a three step pipeline: first, the latex source files for openly licensed papers were downloaded from ArXiv’s bulk access S3 bucket; next, the LATEXML conversion tool was used to convert these source files into a single HTML document; finally, the HTML was converted to plaintext using the Trafilatura HTML processing library. Code for collecting, processing, and preparing this dataset is available in the common pile GitHub repo. Dataset Statistics Documents UTF 8 GB 321,336 21 License Issues While we aim to produce datasets with completely accurate licensing information, license laundering and inaccurate metadata can cause us to erroneously assign the incorrect license to some documents (for further discussion of this limitation, please see our paper). If you believe you have found an instance of incorrect lic…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy