GneissWeb Annotations GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus. This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high quality corpora for research projects, language models, and specialized applications. Learn more about the annotation process and methodology in our official blog post. What's Inside GneissWeb Annotations uses the GneissWeb bloom filter made publicly available by IBM, along with IBM’s Data Prep Kit (now a Linux Foundation AI & Data project) and the GneissWeb groups’ category classifiers. Medical Health information, medical research, and clinical content Education Learning materials, academic resources, and educational platforms Technology Software documentation, technical guides, and tech industry content Science Research publications, scientific articles, and academic work You can access annotations at two levels of granularity: URL level Individual URL classifications for precise content selection (This dataset) Host level Aggregate statistics for entire do…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy