Llama 4 Scout 17B 16E Instruct FP8 dynamic Model Overview Model Architecture: Llama4ForConditionalGeneration Input: Text / Image Output: Text Model Optimizations: Activation quantization: FP8 Weight quantization: FP8 Release Date: 04/15/2025 Version: 1.0 Validated on: RHOAI 2.20, RHAIIS 3.0, RHELAI 1.5 Model Developers: Red Hat (Neural Magic) Model Optimizations This model was obtained by quantizing activations and weights of Llama 4 Scout 17B 16E Instruct to FP8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix multiply compute throughput (by approximately 2x). Weight quantization also reduces disk size requirements by approximately 50%. The llm compressor library is used for quantization. Deployment This model can be deployed efficiently on vLLM, Red Hat Enterprise Linux AI, and Openshift AI, as shown in the example below. Deploy on vLLM vLLM also supports OpenAI compatible serving. See the documentation for more details. Deploy on Red Hat AI Inference Server Deploy on Red Hat Enterprise Linux AI See Red Hat Enterprise Linux AI documentation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy