gemma 4 E4B it Q4 K M Distributed GGUF inference package for Mesh LLM GGUF layer package for running gemma 4 E4B it Q4 K M across a local Mesh LLM cluster. This package is derived from unsloth/gemma 4 E4B it GGUF and keeps the original GGUF distribution split into per layer artifacts for distributed inference. Highlights Run locally Pool multiple machines OpenAI compatible Package variant Private inference on your hardware Split layers across peers Serve /v1/chat/completions locally Q4 K M layer package Model Overview Property Value Source model unsloth/gemma 4 E4B it GGUF Model id unsloth/gemma 4 E4B it GGUF:Q4 K M Family Gemma Parameter scale 4B Quantization Q4 K M Layer count 42 Activation width 2560 Package size 7.5 GB Source file gemma 4 E4B it Q4 K M.gguf Package repo meshllm/gemma 4 E4B it Q4 K M layers Recommended Use Local and private inference with Mesh LLM. Multi machine serving when the full GGUF is too large for one host. OpenAI compatible chat/completions workflows through Mesh LLM's local API. For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmark notes, see the source model card: unsloth/gemma 4 E4B it GGUF.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy