Zyda Zyda is a 1.3T language modeling dataset created by collecting open and high quality datasets and combining them and performing a uniform filtering and deduplication step. We find that Zyda performs extremely well in ablations and is at least comparable and potentially better to the best openly available datasets available, due to our meticulous post processing pipeline. We think the best use of Zyda is either as a standalone dataset for language model training up to the 1T scale, or in combination with Fineweb or Dolma for multi trillion token training. An early version of Zyda was used as the primary dataset for phase 1 pretraining of Zamba, a model which performs strongly on a per token basis, testifying to the strength of Zyda as a pretraining dataset. Models trained on Zyda significantly outperform identical models of the Pythia suite trained on the Pile for 300B tokens. Zyda also outperforms Dolma, RefinedWeb, and Fineweb on 1.4B models trained on 50B tokens of each dataset. According to our evaluations, Zyda is the most performant per token open dataset available in its non starcoder variant on language tasks. The Zyda starcoder variant ties with fineweb. These results…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy