Sampled version of cerebras/SlimPajama 627B. Since the original data was shuffled before chunking, I only downloaded train/chunk1 (of 10 total) and further sampled 10%. This should result in roughly 6B tokens, hence SlimPajama 6B. The dataset is 24GBs in storage size when decompressed (original dataset is over 2TBs) and has 5489000 rows. The validation set and test set were sampled as well. Data source proportions for SlimPajama 627B and SlimPajama 6B For sanity purpose, I caluclated the byte proportion of the sampled version. Data source SlimPajama 627B SlimPajama 6B Commoncrawl 52.2% 54.1% C4 26.7% 28.7% GitHub 5.2% 4.2% Books 4.2% 3.7% ArXiv 4.6% 3.4% Wikpedia 3.8% 3.1% StackExchange 3.3% 2.8% Please refer to the original dataset for other info.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy