gemma 4 31B it qat UD Q4 K XL Distributed GGUF inference package for Mesh LLM GGUF layer package for running gemma 4 31B it qat UD Q4 K XL across a local Mesh LLM cluster. This package is derived from unsloth/gemma 4 31B it qat GGUF and keeps the original GGUF distribution split into per layer artifacts for distributed inference. Highlights Run locally Pool multiple machines OpenAI compatible Package variant Private inference on your hardware Split layers across peers Serve /v1/chat/completions locally UD Q4 K XL layer package Model Overview Property Value Source model unsloth/gemma 4 31B it qat GGUF Model id unsloth/gemma 4 31B it qat GGUF:UD Q4 K XL Family Gemma Parameter scale 31B Quantization UD Q4 K XL Layer count 60 Activation width 5376 Package size 17.0 GB Source file gemma 4 31B it qat UD Q4 K XL.gguf Package repo meshllm/gemma 4 31B it qat UD Q4 K XL layers Recommended Use Local and private inference with Mesh LLM. Multi machine serving when the full GGUF is too large for one host. OpenAI compatible chat/completions workflows through Mesh LLM's local API. For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmark note…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy