🏀nlySports Dataset
Overview
OnlySports Dataset is a comprehensive collection of English sports documents, comprising a diverse range of content including news articles, blogs, match reports, interviews, and tutorials. This dataset is part of the larger OnlySports collection, which includes:
- OnlySportsLM: A 196M parameter sports-domain language model
- OnlySports Dataset: The dataset described in this README
- OnlySports Benchmark: A novel evaluation method for assessing sports knowledge generation
Dataset Specifications
- Size: 1.2 TB disks space
- Token Count: Approximately 600 billion RWKV/GPT2 tokens
- Time Span: 2013 to present
- Source: Extracted from FineWeb dataset, a cleaned and deduplicated subset of CommonCrawl
Data Pipeline
The creation of the OnlySports Dataset involved a two-step process:
-
URL Filtering:
-
Applied the following list of sports-related keywords to URLs in FineWeb: football, soccer, basketball, baseball, tennis, athlete, running, marathon, copa, new, nike, adidas, cricket, rugby, golf, volleyball, sports, sport, Sport, wrestling, wwe, hockey, volleyball, cycling, swim, athletic, league, team, champion, playoff, olympic, premierleague, laliga, bundesliga, seriea, ligue1, epl, racing, nascar, motogp, cup, worldcup, fitness, workout, gym, nfl, nba, NBA, NFL, MLB, NHL, FIFA, UEFA, NCAA, MMA, UFC, ufc, mlb, nhl, fifa, uefa, ncaa, boxing, espn, bleacherreport, mma, sicom, formula1, f1, goal.
-
This step reduced the dataset size by approximately 85%
-
-
Custom Sports Text Classifier:
- Developed a specialized classifier to accurately identify and extract sports-related documents
- Based on the Snowflake-arctic-embed-xs model with an added binary classification layer
- Achieved 99% accuracy in distinguishing between sports and non-sports documents
Significance
The OnlySports Dataset represents a major advancement in sports-related text data:
- Largest sport domain dataset to date
- Significantly surpasses previous collections in both scale and comprehensiveness
- Offers researchers and developers an unprecedented resource for training language models and conducting sports-related NLP tasks
Usage and Applications
The OnlySports Dataset can be used for various purposes, including:
- Training domain-specific language models for sports
- Conducting research on sports-related natural language processing tasks
- Developing applications for sports content analysis and generation
OnlySportsLM
As part of the OnlySports collection, the OnlySportsLM was trained on this dataset. Key features of the model include:
- 196M parameters
- Based on the RWKV-v6 architecture
- 20-layer, 640-dimension structure
- Trained on approximately half of the OnlySports Dataset (315B tokens)
For more information on the model and its performance, please refer to Hugginface page
Citation
If you use the OnlySports Dataset in your research, please cite our paper.
Contact
For more information or inquiries about the OnlySports Dataset, please visit our GitHub repository.