S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc full.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy