Quants of https://huggingface.co/coder3101/gemma 4 26B A4B it heretic, using Unsloth's imatrix. Not sure how many different quants I can upload as I am severely constrained by my PC's storage space. Usage Google recommends the following sampler settings: For creative writing, I use: For the image encoder: While it's reasoning is not excessive, sometimes I do want to limit it: Within llama.cpp and koboldcpp, ensure that swa full is enabled as this model uses Sliding Window Attention (SWA). Thinking In order to enable thinking, add at the top of your system prompt: Conversely, to disable thinking simply omit from your system prompt: To parse the thinking, use the following in SillyTavern or your platform of choice: prefix: thought postfix: Reproduction You can read REPRODUCE.md in the repo's "files and versions" to see how I made the quants. mmproj files are also located there. FAQ What quant should I use? The largest model that fits inside your VRAM. Context GBs: context / 8k = VRAM usage. So 16384 / 8192 = 2GB As an example to estimate VRAM cost: In case you want to mainly do OCR tasks, prefer a lower text model quant and a higher mmproj quant (bf16/f16/f32) as encoders are far mor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy