Meta Llama 3.1 8B FP8 Model Overview Model Architecture: Meta Llama 3.1 Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Intended Use Cases: Intended for commercial and research use in multiple languages. Similarly to Meta Llama 3.1 8B, this model serves as a base version. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 7/23/2024 Version: 1.0 License(s): llama3.1 Model Developers: Neural Magic Quantized version of Meta Llama 3.1 8B. It achieves an average score of 65.90 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 66.47. Model Optimizations This model was obtained by quantizing the weights and activations of Meta Llama 3.1 8B to FP8 data type, ready for inference with vLLM built from source. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per tensor quantization is applied, in which a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy