Meta Llama 3 8B Instruct FP8 Model Overview Model Architecture: Meta Llama 3 Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta Llama 3 8B Instruct, this models is intended for assistant like chat. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 6/8/2024 Version: 1.0 License(s): Llama3 Model Developers: Neural Magic Quantized version of Meta Llama 3 8B Instruct. It achieves an average score of 68.22 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 68.71. Model Optimizations This model was obtained by quantizing the weights and activations of Meta Llama 3 8B Instruct to FP8 data type, ready for inference with vLLM = 0.5.0. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per tensor quantization is appli…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy