See our collection for versions of Deepseek R1 including GGUF & 4 bit formats. Unsloth's DeepSeek R1 1.58 bit + 2 bit Dynamic Quants is selectively quantized, greatly improving accuracy over standard 1 bit/2 bit. Instructions to run this model in llama.cpp: Or you can view more detailed instructions here: unsloth.ai/blog/deepseekr1 dynamic 1. Do not forget about and tokens! Or use a chat template formatter 2. Obtain the latest llama.cpp at https://github.com/ggerganov/llama.cpp. You can follow the build instructions below as well: 3. It's best to use min p 0.05 to counteract very rare token predictions I found this to work well especially for the 1.58bit model. 4. Download the model via: 5. Example with Q4 0 K quantized cache Notice no cnv disables auto conversation mode Example output: 6. If you have a GPU (RTX 4090 for example) with 24GB, you can offload multiple layers to the GPU for faster processing. If you have multiple GPUs, you can probably offload more layers. 7. If you want to merge the weights together, use this script: MoE Bits Type Disk Size Accuracy Link Details 1.58bit UD IQ1 S 131GB Fair Link MoE all 1.56bit. down proj in MoE mixture of 2.06/1.56bit 1.73bit UD IQ1 M…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy