Grok 1 GGUF Quantizations This repository contains unofficial GGUF Quantizations of Grok 1, compatible with llama.cpp as of PR Add grok 1 support 6204. Updates Native Split Support in llama.cpp The splits have been updated to utilize the improvements from PR: llama model loader: support multiple split/shard GGUFs. As a result, manual merging with gguf split is no longer required. With this, there is no need to merge the split files before use. Just download all splits and run llama.cpp with the first split like you would previously. It'll detect the other splits and load them as well. Direct Split Download from huggingface using llama.cpp Thanks to a new PR common: llama load model from url split support 6192 from phymbert it's now possible load model splits from url. That means this downloads and runs the model: And that is very cool (@phymbert) Available Quantizations The following Quantizations are currently available for download: Quant Split Files Size Q2 K 1 of 9, 2 of 9, 3 of 9, 4 of 9, 5 of 9, 6 of 9, 7 of 9, 8 of 9, 9 of 9 112.4 GB IQ3 XS 1 of 9, 2 of 9, 3 of 9, 4 of 9, 5 of 9, 6 of 9, 7 of 9, 8 of 9, 9 of 9 125.4 GB Q4 K 1 of 9, 2 of 9, 3 of 9, 4 of 9, 5 of 9, 6 of 9, 7 o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy