Zyda 2 Zyda 2 is a 5 trillion token language modeling dataset created by collecting open and high quality datasets and combining them and cross deduplication and model based quality filtering. Zyda 2 comprises diverse sources of web data, highly educational content, math, code, and scientific papers. To construct Zyda 2, we took the best open source datasets available: Zyda, FineWeb, DCLM, and Dolma. Models trained on Zyda 2 significantly outperform identical models trained on the Pile, RefinedWeb, FineWeb, FineWeb Edu, and DCLM. Due to our post processing deduplication, filtering, and weighting pipeline, Zyda 2 outperforms all its constituent datasets in resulting model quality. An early version of Zyda 2 was used as the primary dataset for phase 1 pretraining of our Zamba2 series of models which perform extremely strongly on a per token basis and are often state of the art for their size, testifying to the strength of Zyda 2 as a pretraining dataset. According to our evaluations, Zyda 2 is the most performant per token open dataset available. Zyda 2 excels at educational and natural language reasoning content. For code performance, we recommend mixing it with a pure code dataset…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy