Bespoke Stratos 17k We replicated and improved the Berkeley Sky T1 data pipeline using SFT distillation data from DeepSeek R1 to create Bespoke Stratos 17k a reasoning dataset of questions, reasoning traces, and answers. This data was used to train: 1. Bespoke Stratos 32B, a 32B reasoning model which is a fine tune of Qwen 2.5 32B Instruct 2. Bespoke Stratos 7B, a 7B reasoning model which is a fine tune of Qwen 2.5 7B Instruct. Metrics for Bespoke Stratos 32B Metric Bespoke Stratos 32B Sky T1 32B o1 preview DeepSeek R1 DeepSeek R1 Distill Qwen 32B (Ours) DeepSeek R1 Distill Qwen 32B (Reported) AIME2024 63.3 43.3 40.0 79.8 66.7 72.6 MATH500 93.0 82.4 81.4 97.3 89.8 94.3 GPQA Diamond 58.1 56.8 75.2 71.5 61.1 62.1 LCB v2 Easy 96.7 86.3 92.9 91.2 LCB v2 Medium 75.2 56.8 54.9 75.7 LCB v2 Hard 26.2 17.9 16.3 38.2 LCB v2 All 71.1 57.9 59.1 72.2 Metrics for Bespoke Stratos 7B Bespoke Stratos 7B Qwen2.5 7B Instruct DeepSeek R1 Distill Qwen 7B (Ours) DeepSeek R1 Distill Qwen 7B (Reported) AIME2024 20.0 10.0 43.3 55.5 MATH500 82.0 74.2 89.4 92.8 GPQA Diamond 37.8 33.3 44.9 49.1 LiveCodeBench v2 Easy 71.4 65.9 81.3 LiveCodeBench v2 Medium 25.5 18.9 42.2 LiveCodeBench v2 Hard 1.6 3.3 2.4 LiveCo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy