π MINT 1T: Scaling Open Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens π MINT 1T is an open source M ultimodal INT erleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale up from existing open source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. π MINT 1T is designed to facilitate research in multimodal pretraining. π MINT 1T is created by a team from the University of Washington in collaboration with Salesforce Research, other academic institutions including Stanford University, University of Texas at Austin, and University of California Berkeley. You are currently viewing the ArXiv subset of π MINT 1T. For HTML and PDF subsets, please refer to the π MINT 1T collection. Dataset Details Dataset Sources Repository : https://github.com/mlfoundations/MINT 1T Paper: https://arxiv.org/abs/2406.11271 Blog: https://blog.salesforceairesearch.com/mint 1t/ Uses Direct Use π MINT 1T is designed to facilitate research in multimodal pretraining. The dataset can be used for training multimodal models that can reson about interleaved text and images sequences such as Idefics2, XGen MM, anβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy