Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE toolkit. Files are grouped as / /part .parquet . The Hugging Face dataset viewer uses one config per source dataset and one split safe language name per language. Each row contains dataset , record id , opening comment , language , path , repo , extracted at , metadata , comment license detection , and comment license score . For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable schema. Card updated with the stack v2 dedup export at 2026 07 03T09:31:17+00:00. License Detection comment license detection is a JSON string produced by ScanCode Toolkit. It records whether the opening comment matched a known license notice, the detected license expressions, filtered license matches, ScanCode scan errors, and the best raw ScanCode license score. comment license score is a numeric column with the best raw ScanCode license match score for every row, or 0 when no license match was found. The contains license notice flag is counted with minimum ScanCode score 95 and minimum match coverage 95.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy