roleplay sillytavern llama3 My GGUF IQ Imatrix quants for Sao10K/L3 8B Stheno v3.2. Sao10K with Stheno again, another banger! I recommend checking his page for feedback and support. [!IMPORTANT] Quantization process: For future reference, these quants have been done after the fixes from 6920 have been merged. Imatrix data was generated from the FP16 GGUF and conversions directly from the BF16 GGUF. This was a bit more disk and compute intensive but hopefully avoided any losses during conversion. If you noticed any issues let me know in the discussions. [!NOTE] General usage: Use the latest version of KoboldCpp . For 8GB VRAM GPUs, I recommend the Q4 K M imat (4.89 BPW) quant for up to 12288 context sizes. Presets: Some compatible SillyTavern presets can be found here (Virt's Roleplay Presets) . Check discussions such as this one for other recommendations and samplers. [!TIP] Personal support: I apologize for disrupting your experience. Currently I'm working on moving for a better internet provider. If you want and you are able to ... You can spare some change over here (Ko fi) . Author support: You can support the author at their own page . Click here for the original model card in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy