Felix Midjourney Archive A deduplicated, checksum addressed preservation dataset of AI generated images created by Felix / waffles13 with Midjourney. Images are stored in deterministic WebDataset TAR shards with searchable Parquet, JSONL, CSV, and SQLite catalogs. Contents Unique images: 32,477 Exact duplicate source copies excluded: 12,560 Images with full embedded prompts: 13,914 Unique Midjourney Job IDs represented: 22,859 Total image bytes before TAR overhead: 46.83 GiB Each sample uses a stable content derived ID ( mj plus the first 24 hexadecimal characters of its SHA 256). Every image has a same key .json sidecar inside its WebDataset shard. metadata.parquet is the primary fast search catalog. Provenance The sources are local Midjourney downloads, historical Midjourney ZIP exports, and older organized folders belonging to the creator. Exact duplicates are identified by full file SHA 256, not by filenames or perceptual similarity. Distinct encodings, crops, resolutions, variants, and upscales remain distinct. Prompts are recovered in this order: 1. embedded PNG Description metadata; 2. creator maintained catalog metadata; 3. cleaned historical filename text. The catalog reco…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy