DCVLM Pool (small) — per sample annotations Every filtering annotation we computed for the small data pool of our DataComp VLM paper: image quality, image–text alignment, language ID, text quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets . These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline consumes to build a training mixture: you pick fields, pick thresholds, and get a subset. Releasing them means you can reproduce our filters, or design your own, without recomputing anything on 828 GB of raw data . The pool these annotate is mlfoundations/dcvlm pool small — 12,206 WebDataset shards, 10.3 TB. Layout One JSON file per pool shard, under the source dataset's directory — exactly mirroring the pool's own / .tar layout , so ai2d/000000.json annotates ai2d/000000.tar . Each file is a JSON object mapping the sample's WebDataset key → its annotations. For example: The annotations Family Fields What it is : : Image–text alignment 48 CLIP style scores from 4 encoders ( clip vitl14 224 , dfn public , siglip…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy