text to image 2M: A High Quality, Diverse Text to Image Training Dataset Overview text to image 2M is a curated text image pair dataset designed for fine tuning text to image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text to image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better fine tuning results. However, existing publicly available datasets often have limitations: Image Understanding Datasets : Not guarantee the quality of image. Informal collected or Task Specific Datasets : Not category balanced or lacks diversity. Size Constraints : Available datasets are either too small or too large. (subset sampled from large datasets often lack diversity.) To address these issues, we combined and enhanced existing high quality datasets using state of the art text to image and captioning models to create text to image 2M . This includes data 512 2M, a 2M 512x512 fine tuning dataset and data 1024 10K, a 10K high quality, high resolution dataset (for high resolution adaptation). Dataset Composition data 512…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy