Fine T2I: An Open, Large Scale, and Diverse Dataset for High Quality T2I Fine Tuning [[arxiv]](https://arxiv.org/abs/2602.09439) by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [[Dataset Explore]](https://huggingface.co/spaces/ma xu/fine t2i explore) to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20] : Fine T2I reaches the 1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16] : Fine T2I tops the Hugging Face Datasets Trending list, reaching the 2 spot and 1 spot for image datasets! ⭐️⭐️⭐️ Fine T2I is a large scale, high quality, and fully open dataset designed to advance SOTA text to image (T2I) fine tuning. Comprising over 6 million text–image pairs (approximately 2 TB) , Fine T2I was constructed to bridge the performance gap between open community models and enterprise grade models. The dataset distinguishes itself through a rigorous construction pipeline that combines high fidelity synthetic data with professional real world photography, ensuring exceptional visual quality and precise instruction adherence. Key Features Massive Scale & Quality: Contains ~6.15M synthetic samples generated by SOTA dif…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy