This repo contains the EMB BF16 quants, which are quantizations that preserve full precision for embeddings that always stay in system RAM. These embeddings are never copied to GPU VRAM, so they do not contribute to VRAM usage. They're lookup tables that are only accessed a few KB per token so they don't affect RAM bandwidth or computation efficiency despite their larger unquantized size. If you're not constrained by system RAM, having these embeddings in full precision may provide better accuracy. These embeddings are: token embd.weight position embd.weight token types.weight per layer token embd.weight For this model, token embd.weight is left unquantized in the EMB BF16 quants. 🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨 I can no longer upload new models unless I can cover the cost of additional storage. I host 70+ free models as an independent contributor and this work is unpaid. Without your support, no more new models can be uploaded. 🎉 Patreon (Monthly) ☕ Ko fi (One time) Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone. This is the full model with the 19 MTPs all intacts. 88% fewer refusals (10/…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy