π MINT 1T: Scaling Open Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens π MINT 1T is an open source M ultimodal INT erleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale up from existing open source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. π MINT 1T is designed to facilitate research in multimodal pretraining. π MINT 1T is created by a team from the University of Washington in collaboration with Salesforce Research, other academic institutions including Stanford University, University of Texas at Austin, and University of California Berkeley. You are currently viewing the HTML subset of π MINT 1T. For PDF and ArXiv subsets, please refer to the π MINT 1T collection. Updates 9/7/24 We have improved MINT 1T (HTML) by removing boilerplate from the header and footer of each document. This new version of the data can be found in directory data v1 1 and contains 742B text tokens. The previous version of the data can be found in directory data v1 0 . 8/8/24 We have updated MINT 1T (HTML) with fixed document URL filtering and additional image safety filtering. As we prioritizeβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy