Dataset Card for ArxivCap Table of Contents Dataset Card for ArxivCap Table of Contents Dataset Description Dataset Summary Curation Process Dataset Structure Data Loading Data Fields Data Instances Additional Information Licensing Information Citation Information Dataset Description Paper: Multimodal ArXiv Point of Contact: nlp.lilei@gmail.com HomePage : https://mm arxiv.github.io/ Data Instances Example 1 of single (image, caption) pairs "......" stands for omitted parts. Example 2 of multiple images and subcaptions "......" stands for omitted parts. Dataset Summary The ArxivCap dataset consists of 6.4 million images and 3.9 million captions with 193 million words from 570k academic papers accompanied with abstracts and titles. (papers before June 2023 ) Curation Process Refer to our paper for the curation and filter process. Dataset Structure Data Loading Data Fields One record refers to one paper: src: String . "\ /\ "e.g. "arXiv src 2112 060/2112.08947" arxiv id: String . Arxiv id of the paper, e.g. "2112.08947" title: String . Title of the paper. abstract: String . Abstract of the paper. meta: meta from kaggle: refers to arXiv Dataset journey: String . Information about the j…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy