mlabonne/Meta Llama 3.1 8B Instruct abliterated AWQ Model creator: mlabonne Original model: Meta Llama 3.1 8B Instruct abliterated How to use Install the necessary packages Example Python code About AWQ AWQ is an efficient, accurate and blazing fast low bit weight quantization method, currently supporting 4 bit quantization. Compared to GPTQ, it offers faster Transformers based inference with equivalent or better quality compared to the most commonly used GPTQ settings. AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead. It is supported by: Text Generation Webui using Loader: AutoAWQ vLLM version 0.2.2 or later for support for all model types. Hugging Face Text Generation Inference (TGI) Transformers version 4.35.0 and later, from any code or client that supports Transformers AutoAWQ for use from Python code
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy