Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE toolkit. Files are grouped as / /part .parquet . The Hugging Face dataset viewer uses one config per source dataset and one split safe language name per language. Each row contains dataset , record id , opening comment , language , path , repo , extracted at , metadata , comment license detection , and comment license score . For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable schema. Card updated with the stack v2 dedup export at 2026 07 03T09:31:17+00:00. License Detection comment license detection is a JSON string produced by ScanCode Toolkit. It records whether the opening comment matched a known license notice, the detected license expressions, filtered license matches, ScanCode scan errors, and the best raw ScanCode license score. comment license score is a numeric column with the best raw ScanCode license match score for every row, or 0 when no license match was found. The contains license notice flag is counted with minimum ScanCode score 95 and minimum match coverage 95.
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy