SwallowCode v2 Resources 📑 arXiv : Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset : Discover SwallowMath v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode v1 was a high quality Python code dataset generated through an LLM based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License , and (2) its size was limited to 16.1 B tokens, restricting large scale pre training. To address these issues, we built SwallowCode v2 , a fully rewritten Python corpus derived from The Stack v2 , using Qwen3 235B A22B Instruct. The resulting dataset contains 49.8 billion tokens and is released under the Apache 2.0 License , ensuring both open accessibility and reproducibility for research and commercial use. As shown in the figure below, SwallowCode v2 demonstrates stronger performance than other open source code datasets on downstream code generation benchmarks. † Note: While datasets such as OpenCoder and NVIDIA/Nemotron Pretraining Code v1 are labeled “open,” they only release metadata, not the actual training samples. Unlike The Stack v2, they cannot be…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy