Pretraining Effectively on S2ORC ! The peS2o dataset is a collection of ~40M creative open access academic papers, cleaned, filtered, and formatted for pre training of language models. It is derived from the [Semantic Scholar Open Research Corpus]2, or S2ORC. We release multiple version of peS2o, each with different processing and knowledge cutoff date. We recommend you to use the latest version available. If you use this dataset, please cite: Document Format Each document in the dataset is a dictionary with the following fields: added : Date the document was added to the corpus. created : Best guess date for when the document was first published. Some have resolution down to the day, only down to the year. id : Semantic Scholar Corpus ID of the document; it can be used with the Semantic Scholar API to retrieve metadata about the document (e.g., fields of study, authors). source : Collection from which the document was sourced from. At the moment, two are supported: s2orc : collection of full text papers s2ag : collection of title and abstracts text : Text of the document. Paragraphs are separated by two newlines ( \n\n ). version : version of peS2o. peS2o V2 (Latest) Key Facts Kno…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy