🏀nlySports Dataset Overview OnlySports Dataset is a comprehensive collection of English sports documents, comprising a diverse range of content including news articles, blogs, match reports, interviews, and tutorials. This dataset is part of the larger OnlySports collection, which includes: 1. OnlySportsLM: A 196M parameter sports domain language model 2. OnlySports Dataset: The dataset described in this README 3. OnlySports Benchmark: A novel evaluation method for assessing sports knowledge generation Dataset Specifications Size: 1.2 TB disks space Token Count: Approximately 600 billion RWKV/GPT2 tokens Time Span: 2013 to present Source: Extracted from FineWeb dataset, a cleaned and deduplicated subset of CommonCrawl Data Pipeline The creation of the OnlySports Dataset involved a two step process: 1. URL Filtering: Applied the following list of sports related keywords to URLs in FineWeb: football, soccer, basketball, baseball, tennis, athlete, running, marathon, copa, new, nike, adidas, cricket, rugby, golf, volleyball, sports, sport, Sport, wrestling, wwe, hockey, volleyball, cycling, swim, athletic, league, team, champion, playoff, olympic, premierleague, laliga, bundesliga, se…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy