Dataset Card for "ArtifactAI/arxiv s2orc parsed" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv s2orc parsed Dataset Summary AlgorithmicResearchGroup/arxiv s2orc parsed is a subset of the AllenAI S2ORC dataset, a general purpose corpus for NLP and text mining research over scientific papers, The dataset is filtered strictly for ArXiv papers, including the full text for each paper. Github links have been extracted from each paper to aid in the development of AlgorithmicResearchGroup/arxiv python research code How to use it Dataset Structure Data Instances Each data instance corresponds to one file. The content of the file is in the text feature, and other features provide some metadata. Data Fields title (sequence): list of titles. author (sequence): list of authors. authoraffiliation (sequence): list of institution affiliations for each author. venue : (integer): paper publication venue. doi : (float): paper doi. pdfurls : (integer): url link to the paper. corpusid : (int): corpus ID as defined by s2orc. arxivid : (int): arxiv paper id. pdfsha : (string): unique pdf hash. text : (string): full text of the arxiv paper. github urls: (sequence): lis…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy