Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards Table of Contents What is it? What is being released? Breakdown by type and year How to download and use this dataset Dataset statistics Monthly breakdown Pipeline status Dataset card Dataset summary Dataset structure Dataset creation Considerations for using the data Additional information What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005 12 through 2026 02 . Right now the archive has 12.1B items (9.7B comments, 2.4B submissions) in 1.1 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards you can load individually or stream together. Reddit has been around since 2005. Millions of people use it to talk about everything programming, sports, cooking, politics, niche hobbies. That makes it one of the best sources of natural conversation data for language model training, sentiment analysis, community research, and information retrieval. Most Reddit datasets only cover specific subreddits or time windows. This one covers al…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy