DCLM baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm baseline 1.0, where all the files have been mapped to a parquet format. DCLM baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral 0.3 7B ? ✗ 57.0 62.7 45.1 QWEN 2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗ 57.6 66.2 46.3 Gemma 8B 6T ✗ 57.8 64.3 44.6 Phi 3 7B ? ✗ 61.0 69.9 57.9 Open weights, open datasets Falcon 7B 1T ✓ 44.1 27.4 25.1 Amber 7B 1.2T ✓ 39.8 27.9 22.3 Crystal 7B 1.2T ✓ 48.0 48.2 33.2 OLMo 1.7 7B 2.1T ✓ 47.0 54.0 34.2 MAP Neo 7B 4.5T ✓ 50.2 57.1 40.4 Models we trained FineWeb edu 7B 0.14T ✓ 38.7 26.3 22.1 FineWeb edu 7B 0.28T ✓ 41.9 37.3 24.5 DCLM BASELINE 7B 0.14T ✓ 44.1 38.3 25.0 DCLM BASELINE 7B 0.28T ✓ 48.9 50.8 31.8 DCLM BASELINE 7B 2.6T ✓ 57.1 63.7 45.4 Dataset Details Dataset Description Curated by: The DCLM Team Language(s) (NLP): English License: CC by 4.0 Data…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy