huihui ai/Huihui DeepSeek V4 Flash abliterated ds4 GGUF This is an uncensored version of deepseek ai/DeepSeek V4 Flash created with abliteration. This quants are specific for the DS4(antirez/ds4) and llama.cpp inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader). Note 1. The Q2 version has a certain refusal rate. It should be fine for writing code, while the other versions are still under testing. 2. Choose the appropriate model based on the size of your GPU. All models can run under both Fringe210/llama.cpp deepseek v4 flash cuda (supports multi GPU) and ds4 (supports multi GPU). 3. ds4 now supports multi GPU operation. For more information on how to use it, please refer to x.com/support huihui DS4 Unix Domain Socket (UDS) Acceleration Patch Dramatically accelerate multi GPU layer splitting inference on the same machine (coordinator + worker mode) by replacing TCP loopback with Unix Domain Sockets. open source 👉 huihui support/ds4/tree/uds DS4 Tensor Parallel Acceleration Patch Dramatically speed up multi GPU layer splitting inference on a single machine using a single process, with full support…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy