π MINT 1T: Scaling Open Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens π MINT 1T is an open source M ultimodal INT erleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale up from existing open source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. π MINT 1T is designed to facilitate research in multimodal pretraining. π MINT 1T is created by a team from the University of Washington in collaboration with Salesforce Research, other academic institutions including Stanford University, University of Texas at Austin, and University of California Berkeley. You are currently viewing a subset of the PDF portion of π MINT 1T associated with CommonCrawl dump CC 2023 50 . For other PDF, HTML, and ArXiv subsets, refer to the π MINT 1T collection. Updates 9/19/24 We have removed roughly 10% of the PDF samples as there was a mismatch between the frames in the TIFF images and the document metadata. 8/8/24 We have become aware that the image hashes in the PDF subset of MINT 1T do not match the images in the documents. We want to emphasize that the images for each document are correct, and onlβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy