About weighted/imatrix quants of https://huggingface.co/deepseek ai/deepseek moe 16b chat static quants are available at https://huggingface.co/mradermacher/deepseek moe 16b chat GGUF Usage If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi part files. Provided Quants (sorted by size, not necessarily quality. IQ quants are often preferable over similar sized non IQ quants) Link Type Size/GB Notes : : : : GGUF i1 IQ1 S 5.3 for the desperate GGUF i1 IQ1 M 5.6 mostly desperate GGUF i1 IQ2 XXS 6.0 GGUF i1 IQ2 XS 6.3 GGUF i1 IQ2 S 6.4 GGUF i1 IQ2 M 6.7 GGUF i1 Q2 K 6.8 IQ3 XXS probably better GGUF i1 Q2 K S 6.8 very low quality GGUF i1 IQ3 XXS 7.3 lower quality GGUF i1 IQ3 XS 7.5 GGUF i1 IQ3 S 7.9 beats Q3 K GGUF i1 Q3 K S 7.9 IQ3 XS probably better GGUF i1 IQ3 M 8.0 GGUF i1 Q3 K M 8.6 IQ3 S probably better GGUF i1 Q3 K L 8.9 IQ3 M probably better GGUF i1 IQ4 XS 9.0 GGUF i1 IQ4 NL 9.4 prefer IQ4 XS GGUF i1 Q4 0 9.4 fast, low quality GGUF i1 Q4 K S 10.0 optimal size/speed/quality GGUF i1 Q4 1 10.4 GGUF i1 Q4 K M 11.0 fast, recommended GGUF i1 Q5 K S 11.7 GGUF i1 Q5 K M 12.5 GGUF i1 Q6 K 14.8 practically like…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy