π» EAI Taxonomy Code w/ DCLM π Website π₯οΈ Code π Paper A 564 billion token dataset of high quality code curated from web data using taxonomy based filtering. π― Dataset Overview This dataset is part of the Essential Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional code datasets that require complex domain specific pipelines, our approach leverages a 12 category taxonomy to efficiently identify and extract high quality code data. π‘ EAI Taxonomy Code w/ DCLM (564B tokens): Documents targeting code that exhibit intermediate to advanced reasoning, combined with the DCLM classifier to filter for instruction dense documents. Also includes mathematics content ( 51 Mathematics ) to match the scope of existing code datasets. π Performance Our taxonomy based approach achieves competitive results with significantly less curation effort: Dataset HumanEval+ MBPP+ MMLU CS Curation Complexity DCLM baseline 28.0% 45.5% 32.0% General web filtering OpenCoder FW 26.2% 45.8% 27.7% Complex domain pipeline EAI Taxonomy Code 27.4% 46.6% 29.0% Simple semantic filter EAI Taxonomy Code w/ DCLM 28.7% 45.0% 47.0% +β¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy