Chat & support: TheBloke's Discord server Want to contribute? TheBloke's Patreon page TheBloke's LLM work is generously supported by a grant from andreessen horowitz (a16z) Openhermes 2.5 Mistral 7B AWQ Model creator: Teknium Original model: Openhermes 2.5 Mistral 7B Description This repo contains AWQ model files for Teknium's Openhermes 2.5 Mistral 7B. These files were quantised using hardware kindly provided by Massed Compute. About AWQ AWQ is an efficient, accurate and blazing fast low bit weight quantization method, currently supporting 4 bit quantization. Compared to GPTQ, it offers faster Transformers based inference with equivalent or better quality compared to the most commonly used GPTQ settings. It is supported by: Text Generation Webui using Loader: AutoAWQ vLLM Llama and Mistral models only Hugging Face Text Generation Inference (TGI) AutoAWQ for use from Python code Repositories available AWQ model(s) for GPU inference. GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8 bit GGUF models for CPU+GPU inference Teknium's original unquantised fp16 model in pytorch format, for GPU inference and for further conversions Prompt temp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy