We are excited to release a synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek ai/DeepSeek R1 Distill Llama 70B. While there have been multiple efforts to build open reasoning datasets for math and code tasks, we noticed a lack of large datasets containing reasoning traces for diverse non code/math topics like social and natural sciences, education, creative writing and general conversations, which is why we decided to release this dataset. Note: Please note that in this instance we have not verified the reasoning traces and answers for accuracy. Dataset details: Total number of rows: 22.2 million rows Total number of tokens: 35.8 billion tokens The dataset can be used to fine tune smaller, more efficient models to mimic the reasoning capabilities of larger models like DeepSeek R1 using SFT. Response format: Loading the dataset:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy