Dataset Card for ImageNet D This is a FiftyOne dataset with 4838 samples. Installation If you haven't already, install FiftyOne: Usage Dataset Description ImageNet D is a new benchmark created using diffusion models to generate realistic synthetic images with diverse backgrounds, textures, and materials. The dataset contains 4,835 hard images that cause significant accuracy drops of up to 60% for a range of vision models, including ResNet, ViT, CLIP, LLaVa, and MiniGPT 4. To create ImageNet D, a large pool of synthetic images is generated by combining object categories with various nuisance attributes using Stable Diffusion. The most challenging images that cause shared failures across multiple surrogate models are selected for the final dataset. Human labelling via Amazon Mechanical Turk is used for quality control to ensure the images are valid and high quality. Experiments show that ImageNet D reveals significant robustness gaps in current vision models. The synthetic images transfer well to unseen models, uncovering common failure modes. ImageNet D provides a more diverse and challenging test set than prior synthetic benchmarks like ImageNet C, ImageNet 9, and Stylized ImageNet…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy