Llama 3.3 70B Instruct Q3 K M Distributed GGUF inference package for Mesh LLM GGUF layer package for running Llama 3.3 70B Instruct Q3 K M across a local Mesh LLM cluster. This package is derived from unsloth/Llama 3.3 70B Instruct GGUF and keeps the original GGUF distribution split into per layer artifacts for distributed inference. Highlights Run locally Pool multiple machines OpenAI compatible Package variant Private inference on your hardware Split layers across peers Serve /v1/chat/completions locally Q3 K M layer package Model Overview Property Value Source model unsloth/Llama 3.3 70B Instruct GGUF Model id unsloth/Llama 3.3 70B Instruct GGUF:Q3 K M Family Llama Parameter scale 70B Quantization Q3 K M Layer count 80 Activation width 8192 Package size 32.5 GB Source file Llama 3.3 70B Instruct Q3 K M.gguf Package repo meshllm/Llama 3.3 70B Instruct Q3 K M draft layers Recommended Use Local and private inference with Mesh LLM. Multi machine serving when the full GGUF is too large for one host. OpenAI compatible chat/completions workflows through Mesh LLM's local API. For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmar…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy