I removed all low quality data and uploaded it here This dataset is the entire 21K ImageNet dataset with about 13 million examples and about 19 thousand classes as strings (for some reason it only had ~19K classes instead of 21K) as well as the entire CC12M dataset, recaptioned. If you just want the recaptioned Imagenet dataset, I have that here I obtained the CC12M form others. CC12M is a dataset with 12 million images created in 2021. Unfortunately the downloader provided by Google has many broken links and the download takes forever. However, some people in the community publicized the dataset. The largest of these repos I could find where ach image is full resolution is https://huggingface.co/datasets/lmms lab/LLaVA ReCap CC12M with about 10 million images. The captions are very unnatural for image generation, so I merge this data with the data from https://huggingface.co/datasets/CaptionEmporium/conceptual captions cc12m llavanext on ID which has much better captions. Thanks again for these repos!! For the imagenet dataset, I recaptioned everything using the method below. The images are in PNG format. They can be decoded like in the following example where row["image"] are the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy