Gemma 4 31B JANG 4M CRACK GGUF GGUF quantizations of Gemma 4 31B JANG 4M CRACK for use with llama.cpp, LM Studio, Ollama, and other GGUF compatible inference engines. About the Model Base model: google/gemma 4 31b it Architecture: Gemma 4 Dense Transformer (31B parameters, 60 layers) Features: Hybrid Sliding/Global Attention, Vision + Audio multimodal Modification: CRACK abliteration (refusal removal) + JANG v2 mixed precision quantization Why This Conversion? The original model uses JANG v2 mixed precision MLX quantization (attention 8 bit + MLP 4 bit), which is only compatible with vMLX. Standard tools (llama.cpp, LM Studio, oMLX, mlx lm) cannot load this format due to mixed per layer bit widths. This repository provides standard GGUF quantizations that work everywhere. Conversion Process Note: Since the original was already quantized (avg 5.1 bits), the dequantized f16 intermediate is an approximation. Re quantizing to GGUF introduces minimal additional quality loss since the attention layers were preserved at 8 bit in the original. Available Quantizations File Quant Size Quality Notes gemma 4 31b jang crack Q3 K M.gguf Q3 K M ~14 GB Acceptable Minimum viable quality gemma 4 31b…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy