Huihui Gemma 4 26B A4B IT Abliterated — GGUF Quantizations This repository contains GGUF / llama.cpp quantized builds of: huihui ai/Huihui gemma 4 26B A4B it abliterated These are UD quantizations prepared for efficient local inference with llama.cpp , including support for multimodal image text to text workflows when used with the corresponding mmproj file. Overview This release is designed for users who want to run the Huihui Gemma 4 26B A4B abliterated model locally with reduced VRAM and RAM requirements while preserving as much output quality as possible. The quantization variants use an optimized tensor distribution strategy inspired by Unsloth style mixed quality quantization recipes , balancing model fidelity, speed, and memory efficiency across different hardware targets. Quick Start 1. Download the latest release of llama.cpp . 2. Download your preferred .gguf model file from this repository. 3. For multimodal inference, also download the matching mmproj file. 4. Run the model with llama.cpp using your preferred frontend or CLI. Example: Adjust the model filename and mmproj filename to match the files you downloaded. Which Quant Should I Choose? Choose based on your availa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy