Chat & support: TheBloke's Discord server Want to contribute? TheBloke's Patreon page TheBloke's LLM work is generously supported by a grant from andreessen horowitz (a16z) Llama 2 7B Chat GPTQ Model creator: Meta Llama 2 Original model: Llama 2 7B Chat Description This repo contains GPTQ model files for Meta Llama 2's Llama 2 7B Chat. Multiple GPTQ parameter permutations are provided; see Provided Files below for details of the options provided, their parameters, and the software used to create them. Repositories available AWQ model(s) for GPU inference. GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8 bit GGUF models for CPU+GPU inference Meta Llama 2's original unquantised fp16 model in pytorch format, for GPU inference and for further conversions Prompt template: Llama 2 Chat Provided files and GPTQ parameters Multiple quantisation parameters are provided, to allow you to choose the best one for your hardware and requirements. Each separate quant is in a different branch. See below for instructions on fetching from different branches. All recent GPTQ files are made with AutoGPTQ, and all files in non main branches are made with Au…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy