h2o danube3 4b chat GGUF Model creator: H2O.ai Original model: h2oai/h2o danube3 4b chat Description This repo contains GGUF format model files for h2o danube3 4b chat quantized using llama.cpp framework. Table below summarizes different quantized versions of h2o danube3 4b chat. It shows the trade off between size, speed and quality of the models. Name Quant method Model size MT Bench AVG Perplexity Tokens per second : : : : : : : : : : : h2o danube3 4b chat F16.gguf F16 7.92 GB 6.43 6.17 479 h2o danube3 4b chat Q8 0.gguf Q8 0 4.21 GB 6.49 6.17 725 h2o danube3 4b chat Q6 K.gguf Q6 K 3.25 GB 6.37 6.20 791 h2o danube3 4b chat Q5 K M.gguf Q5 K M 2.81 GB 6.25 6.24 927 h2o danube3 4b chat Q4 K M.gguf Q4 K M 2.39 GB 6.31 6.37 967 h2o danube3 4b chat Q3 K M.gguf Q3 K M 1.94 GB 5.87 6.99 1099 h2o danube3 4b chat Q2 K.gguf Q2 K 1.51 GB 3.71 9.42 1299 Columns in the table are: Name model name and link Quant method quantization method Model size size of the model in gigabytes MT Bench AVG MT Bench benchmark score. The score is from 1 to 10, the higher, the better Perplexity perplexity metric on WikiText 2 dataset. It's reported in a perplexity test from llama.cpp. The lower, the better Token…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy