ArXiv Models Data Code Blog Sample Explorer Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, Sean Welleck The Proof Pile 2 is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the Llemma 7B and Llemma 34B models. It consists of three subsets: arxiv (29B tokens): the ArXiv subset of RedPajama open web math (15B tokens): The OpenWebMath dataset, which contains much of the high quality mathematical text from the internet. algebraic stack (11B tokens): A new dataset of mathematical code, including numerical computing, computer algebra, and formal mathematics. You can download the dataset as follows Schema Each dataset row has the following structure Dataset Contents For detailed documentation of the ArXiv and web subsets, refer to RedPajama and OpenWebMath. The following table enumerates the contents of the AlgebraicStack by programming language. The AlgebraicStack is filtered to only include documents that contain mathematics, as judged by hand crafted, language specific heuristics. Language AlgebraicStack tokens Agda 35.2 M C 25.1 M C++ 954.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy