Llamacpp imatrix Quantizations of dolphin 2.9.2 qwen2 7b Using llama.cpp release b2965 for quantization. Original model: https://huggingface.co/cognitivecomputations/dolphin 2.9.2 qwen2 7b All quants made using imatrix option with dataset from here Prompt format Downloading using huggingface cli First, make sure you have hugginface cli installed: Then, you can target the specific file you want: If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run: You can either specify a new local dir (dolphin 2.9.2 qwen2 7b Q8 0) or download them all in place (./) Which file should I choose? A great write up with charts showing various performances is provided by Artefact2 here The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have. If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1 2GB smaller than your GPU's total VRAM. If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy