◎ Genius Lyrics Dataset Cleaned & Deduplicated 🤗 Hugging Face 🤗 Hugging Face DOI: 10.57967/hf/7978 DOI: 10.57967/hf/7978 revision: 9742989 revision: 9742989 A heavily cleaned, English only, genre filtered subset of the Genius Song Lyrics with Language Information Kaggle dataset. Reduced from 9+ GB to 2.56 GB through language filtering, genre filtering, artifact removal, and deduplication. Optimized for language model fine tuning, lyric generation, and music NLP research. No train/validation/test splits This dataset ships as a single train split. You should create validation and test splits appropriate to downstream task Overview The raw Genius dataset contains millions of song entries across dozens of languages, genres, and quality levels — including non music content like poetry, book excerpts, and miscellaneous text tagged as misc . This cleaned version retains only English language songs from verified music genres , with lyrics scrubbed of Genius UI artifacts, HTML residue, and duplicates. The result is a high signal corpus suitable for causal language model pretraining or supervised fine tuning on lyric generation tasks. Dataset Stats .th { font family: Fira Code, Consolas, m…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy