Dataset Card for DeMix Corpora DeMix π Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre training π€ Dataset: DeMix Corpora π± Github: Demix Dataset Details Dataset Description DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high quality, large scale, and carefully mixed resource that can be directly employed for pre training.β¦ See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix Corpora.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy