obswork/arxiv ai ml 100k A 99,999 paper stratified subset of Rendra8631/arxiv papers at revision a2c6afb51332d2744b46308df6917697582f8cd4 , filtered to the primary subjects cs.AI , cs.CV , cs.LG , and stat.ML . Only papers submitted in 2023 2025 are included. This dataset is a build artifact of the OCR benchmark in modal model experiments/arxiv mirror/ . It exists so that the downstream benchmark corpus (rasterized page images) is built from a stable, well scoped PDF pool. Contents metadata.parquet 99999 rows, one per paper. Carries the source schema plus derived primary code , submission year , and target filename columns. pdfs/ / / .pdf one file per paper, bucketed by primary arXiv subject code (e.g. cs.CV ) and submission year month. The YYMM layer keeps every directory under the HF 10k files per dir cap and gives consumers a natural temporal slice. Filter and sampling Primary subjects kept: cs.AI , cs.CV , cs.LG , stat.ML . Submission year kept: = 2023 (derived from the arxiv id YYMM.NNNNN prefix, falling back to a regex over submission date for any pre 2007 old format IDs). Version dedup: for papers with multiple versions in the source metadata ( ...v1 , ...v2 ), only the late…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy