This repo contains GGUF versions of the meta llama/Meta Llama 3 8B Instruct model. Make AI models cheaper, smaller, faster, and greener! Give a thumbs up if you like this model! Read the documentations to know more here Join the Pruna AI community on Discord here to share feedback/suggestions or get help. Frequently Asked Questions How does the compression work? The model is compressed with GGUF. How does the model quality change? The quality of the model output might vary compared to the base model. What is the model format? We use GGUF format. What calibration data has been used? If needed by the compression method, we used WikiText as the calibration data. How to compress my own models? You can request premium access to more compression methods and tech support for your specific use cases here. Downloading and running the models You can download the individual files from the Files & versions section. Here is a list of the different versions we provide. For more info checkout this chart and this guide: Quant type Description Q5 K M High quality, recommended. Q5 K S High quality, recommended. Q4 K M Good quality, uses about 4.83 bits per weight, recomme…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy