Chat & support: TheBloke's Discord server Want to contribute? TheBloke's Patreon page TheBloke's LLM work is generously supported by a grant from andreessen horowitz (a16z) Synthia V3.0 11B AWQ Model creator: Migel Tissera Original model: Synthia V3.0 11B Description This repo contains AWQ model files for Migel Tissera's Synthia V3.0 11B. These files were quantised using hardware kindly provided by Massed Compute. About AWQ AWQ is an efficient, accurate and blazing fast low bit weight quantization method, currently supporting 4 bit quantization. Compared to GPTQ, it offers faster Transformers based inference with equivalent or better quality compared to the most commonly used GPTQ settings. AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead. It is supported by: Text Generation Webui using Loader: AutoAWQ vLLM version 0.2.2 or later for support for all model types. Hugging Face Text Generation Inference (TGI) Transformers version 4.35.0 and later, from any code or client that supports Transformers AutoAWQ for use from Python code Repositories available AWQ model(s) for GPU inference. GPTQ models for GPU inferen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy