Dataset Card for Recap DataComp 1B Recap DataComp 1B is a large scale image text dataset that has been recaptioned using an advanced LLaVA 1.5 LLaMA3 8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open sourced LLaMA 3, a GPT 4 level LLM. Our recaptioning pipeline is simple: first, we fine tune a LLaMA 3 8B powered LLaVA 1.5 and then employ it to recaption 1.3 billion images from the DataComp 1B dataset. Our empirical results confirm that this enhanced dataset, Recap DataComp 1B, offers substantial benefits in training advanced vision language models. For discriminative models like CLIP, we observe enhanced zero shot performance in cross modal retrieval tasks. For generative models like text to image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Curated by: Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, Yuyin Zhou, Cihang Xie License: cc by 4.0 Dataset Sources…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy