DeepSeek R1 Quantization Details This quantized model was created using AutoAWQ version 0.2.8 with quant config : Paper Link 👁️ 1. Introduction We introduce our first generation reasoning models, DeepSeek R1 Zero and DeepSeek R1. DeepSeek R1 Zero, a model trained via large scale reinforcement learning (RL) without supervised fine tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek R1 Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek R1 Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek R1, which incorporates cold start data before RL. DeepSeek R1 achieves performance comparable to OpenAI o1 across math, code, and reasoning tasks. To support the research community, we have open sourced DeepSeek R1 Zero, DeepSeek R1, and six dense models distilled from DeepSeek R1 based on Llama and Qwen. DeepSeek R1 Distill Qwen 32B outperforms OpenAI o1 mini across various benchmarks, achieving new state of the art results for dense models. 2. Model Summary Post Trai…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy