See our collection for versions of Deepseek R1 including GGUF & 4 bit formats. Unsloth's DeepSeek R1 1.58 bit + 2 bit Dynamic Quants is selectively quantized, greatly improving accuracy over standard 1 bit/2 bit. Instructions to run this model in llama.cpp: You can view more detailed instructions in our blog: unsloth.ai/blog/deepseek r1 1. Do not forget about and tokens! Or use a chat template formatter 2. Obtain the latest llama.cpp at https://github.com/ggerganov/llama.cpp 3. Example with Q8 0 K quantized cache Notice no cnv disables auto conversation mode Example output: 4. If you have a GPU (RTX 4090 for example) with 24GB, you can offload multiple layers to the GPU for faster processing. If you have multiple GPUs, you can probably offload more layers. Finetune your own Reasoning model like R1 with Unsloth! We have a free Google Colab notebook for turning Llama 3.1 (8B) into a reasoning model: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1 (8B) GRPO.ipynb ✨ Finetune for Free All notebooks are beginner friendly ! Add your dataset, click "Run All", and you'll get a 2x faster finetuned model which can be exported to GGUF, vLLM or uploaded to Hug…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy