Chess Pre to Post — Pretraining Corpus v1 (20B) Raw tokenized pretraining data for the Chess Pre to Post project, stored as sharded NumPy arrays ( shard XXXX/raw.NNNN.npy ). [!IMPORTANT] This is an earlier, smaller (20B) snapshot and is no longer maintained. The maintained version of this dataset is pavelslab nyu/pretrain v1 54B . Please use that version for any new work — it supersedes this one. Maintained version ➡️ pavelslab nyu/pretrain v1 54B Format Sharded raw token arrays: shard 0000/ … shard 0005/ , each containing raw.NNNN.npy files of tokenized text. Load shards with numpy.load(...) and concatenate as needed for your data loader. Citation If you use this dataset, please cite the maintained version:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy