GLM 4.7 Flash Uncensored Heretic NEO CODE Imatrix MAX GGUF Specialized and Enhanced UNCENSORED/HERETIC GGUF quants for the new GLM 4.7 Flash, 30B A3B MOE, mixture of experts model. [ https://huggingface.co/zai org/GLM 4.7 Flash ] This model can be run on the GPU(s) and/or CPU due to 4 experts activated (appox 2B parameters active). Uncensored / Heretic'ed De censoring by Heretic (special thanks to "Olafangensan") seems to have reduced the size of thinking blocks in some cases and/or "focused" the model more. Default Settings (Most Tasks) temperature: 1.0 top p: 0.95 max new tokens: 131072 REP PEN: 1.1 OR 1.0 (off) (if you get repeat issues) You might also try GLM 4.6 settings (unsloth): temperature = 0.8 top p = 0.6 (recommended) top k = 2 (recommended) max generate tokens = 16,384 That being said, I suggest min context of 8k 16K as final outputs (post thinking) can be long and detailed and in a number of cases has been observed "polishing" the final output one or more times IN the output section. (Model can handle 200k context, non roped.) NON UNCENSORED QUANTS: https://huggingface.co/DavidAU/GLM 4.7 Flash Uncensored Heretic NEO CODE Imatrix MAX GGUF Quants General: Quants and Ima…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy