Gemma 4 12B Gemini 3.5 flash Reasoning Distill GGUF GGUF quantized versions of Ayodele01/Gemma 4 12B Gemini 3.5 flash Reasoning Distill. Model Description This is Google's Gemma 4 12B instruction tuned model, fine tuned on the full 25,000 synthetic reasoning examples dataset WithinUsAI/gemini 3.5 flash distilled 25k using QLoRA via Unsloth. This GGUF model contains quantized versions of the merged model weights. Available Files and Quantizations Filename Quant Type Size Description Gemma 4 12B Gemini 3.5 flash Reasoning Distill bf16.gguf BF16 ~24.4 GB Full precision, best quality Gemma 4 12B Gemini 3.5 flash Reasoning Distill Q8 0.gguf Q8 0 ~12.2 GB High quality, minimal degradation Gemma 4 12B Gemini 3.5 flash Reasoning Distill Q5 K M.gguf Q5 K M ~8.3 GB Balanced (recommended) Gemma 4 12B Gemini 3.5 flash Reasoning Distill Q4 K M.gguf Q4 K M ~7.2 GB Good quality, smaller size Usage with llama.cpp You can run these files using llama.cpp. Prompt Template Gemma 4 chat template format: Training and Distillation Context For details on evaluations, training hyperparameters, and qualitative findings, please refer to the main repository model card: Ayodele01/Gemma 4 12B Gemini 3.5 flash R…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy