UPDATE 2026 07 20: Download & use Google's new chat template from here for better speed and accurancy. See Unsloth's post for details. WARNING: Created with heavy LLM assistance (Zoo Code + DeepSeek V4 Flash). Use at your own discretion. Cosplayed Frankenstein and grafted/"merged" [llmfan46's][QAT BF16] abliterated tensors onto [Unsloth's][QAT UD Q4 K XL] lossless Q4 0 quant. Should yield better accurancy and refusal rate than a naive abliterated Q4 0 quant. This repo contains two variants: "UDmerge Q4 K XL" have the abliterated tensors ( blk.N.attn output.weight , where N is 11 to 22 inclusive) quantized to Q4 0, while "UDmerge Q4 K XXL" quantizes them to Q8 0. The latter improves refusal rate by a lot, while basically not affecting TG speed (your mileage may vary). Use [Unsloth's][QAT UD Q4 K XL] mmproj and mtp GGUF files for multimodal and MTP support. [QAT BF16] [QAT Q4 0] [QAT Q4 K M] QAT UDmerge Q4 K XL QAT UDmerge Q4 K XXL [PTQ BF16] [PTQ Q4 0] [PTQ Q4 K M] Size (GB) 51.6 16.0 16.8 14.2 14.3 51.6 14.5 16.8 PPL 4.945 4.804 5.246 5.236 5.101 5.279 10.17 8.177 KLD 0.0000 0.1158 0.1076 0.0448 0.0291 0.0000 0.7388 0.6130 Refusal 17% 19% 25% 33% 16% 28% 27% 26% MMLU val 80.86% 80.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy