Dataset Card for MONET MONET ( M assive, O pen, N on redundant and E nriched T ext to image dataset) is a large scale, curated image text dataset designed for training text to image (T2I) systems. It contains 104.9 million high quality image text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic ) through successive stages of safety filtering, domain based filtering, exact and near duplicate removal, and re captioning with multiple vision language models, and is further augmented with synthetically generated samples. Each image is released with pre computed embeddings, structured annotations and pre encoded VAE latents to accelerate downstream use. A 4B parameter latent diffusion model trained exclusively on MONET reaches competitive GenEval and DPG scores, demonstrating that MONET lowers the barrier to large scale, reproducible text to image research. Table of Contents Dataset Summary Dataset Sources Curation Pipeline Data Fields Usage Splits Supported Tasks Demos Retrieval & UMAP Building subsets with FAISS indexes Training Biases, Risks, and Limitations Ethical and Responsible Use Maintenance & Contact Changelog Citation 📋…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy