Gemma 4 E4B — Opus Reasoning + Claude Code GGUF GGUF version of our Opus 4.6 reasoning model. Ollama ✅ LM Studio ✅ llama.cpp ✅ Reasoning baked in — no adapter needed. Built by RavenX AI · GGUF converted from MLX source What is this? This is the GGUF version of gemma 4 E4B Agentic Opus Reasoning GeminiCLI mlx 4bit — Gemma 4 E4B with Opus 4.6 reasoning and Claude Code LoRA fused directly into the weights . No adapter needed, no extra config — just load and run with Claude style reasoning baked in. Looking for the Apple Silicon MLX version? → MLX 4 bit model (optimized for Metal GPU) Available Quantizations Quantization Size Use case : : : : Q4 K M 2.7 GB Recommended — best balance of quality and speed Q5 K M 3.1 GB Higher quality, slightly more RAM Q8 0 4.5 GB Near lossless, needs more RAM F16 8.3 GB Full precision GGUF Sizes will be updated once conversion is complete. 🦙 Ollama — One Command With a custom system prompt Create a Modelfile : OpenAI compatible API 💻 LM Studio 1. Open LM Studio 2. Search for deadbydawn101/gemma 4 E4B Agentic Opus Reasoning GeminiCLI GGUF 3. Download any quantization (Q4 K M recommended) 4. Load and chat — reasoning is baked in 🔧 llama.cpp CLI Server…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy