Quantized Model Information [!IMPORTANT] This repository is an AWQ 4 bit quantized version of meta llama/Llama 3.3 70B Instruct , originally released by Meta AI. This model was quantized using AutoAWQ from FP16 down to INT4 using GEMM kernels, with zero point quantization and a group size of 128. Hardware: Intel Xeon CPU E5 2699A v4 @ 2.40GHz, 256GB of RAM, and 2x NVIDIA RTX 3090. Model usage (inference) information for Transformers, AutoAWQ, Text Generation Interface (TGI), and vLLM , as well as quantization reproduction details, are below. Original Model Information The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperform many of the available open source and closed chat models on common industry benchmarks. Model Usage In order to use this quantized model, support is offered for different solutions such as transformers, autoawq, or text generation inference. [!NOTE] In order to run inference with Llama 3.3 70B Instruct AWQ in INT4, around 35 GiB of VRAM are needed for loading the model…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy