We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
S2ORC ArXiv A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata. Dataset Summary Statistic Value Total papers 2,579,762 Total size ~266 GB Format Parquet Split train Dataset Structure Content Fields Field Type Description title string Paper… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_arxiv.
No dataset card provided yet.
Mirrored from an external registry.
Last synced 6/11/2026
Preview not yet available for this dataset