Quants Quant Tamaño Q4 K M 20.5 GiB Q5 K M 23.9 GiB Q6 K 27.4 GiB Q8 0 37.8 GiB auto detect GPU build llamacpp auto.sh Download models Use (llama.cpp optional MTP and image) Note: Adding ' Don't hallucinate.' to the system file greatly improves the quality of the responses " batch size 4096" A higher value means slower decoding, but a lower value means much faster decoding. " batch size 1024" Note 2: This model is more logical than version 1. Although it has MTP, it doesn't reach +250 tokens/s. No changes were applied to the model layers before the conversion to GGUF, only for testing purposes. Although this model was NOT trained, the merge took more than a day, processing approximately +500,000 graphs. With these points in mind, I can improve the next model using dare ties v2, which is not in Mergekit. Note: I think I've messed with Fable now, because if I ask it to do something with MergeKit to improve it, it jumps to a lower tier model. :'( Fable Sometimes he tells me that the 'Qwen' models are boring and that I should work with 'gemma 4'. Just working with Qwen models or variants of the 'Fable' model can lead me down paths that are not correct.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy