S2ORC ArXiv A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata. Dataset Summary Statistic Value Total papers 2,579,762 Total size ~266 GB Format Parquet Split train Dataset Structure Content Fields Field Type Description title string Paper… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc arxiv.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy