About Dataset Dataset name: arXiv academic paper metadata Data source: https://arxiv.org/ Submission date: 1986 04 25 ~ 2025 05 13 (data updated weekly) Number of papers: 2,710,806 (as of 2025.5.14) Fields included: title, author, abstract, journal information, DOI, etc. Data format: json Data volume: 4.58G About ArXiv For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including math, statistics, electrical engineering, quantitative biology, and economics. This rich corpus of information offers significant, but sometimes overwhelming depth. In these times of unique global challenges, efficient extraction of insights from data is essential. To help make the arXiv more accessible, we present a free, open pipeline on Kaggle to the machine readable arXiv dataset: a repository of 1.7 million articles, with relevant features such as article titles, authors, categories, abstracts, full text PDFs, and more. Our hope is to empower new use cases that can lead to the exploration of richer machine learning techniques t…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy