Meta Llama 3 8B Instruct FP8 KV Model Overview Meta Llama 3 8B Instruct quantized to FP8 weights and activations using per tensor quantization, ready for inference with vLLM = 0.5.0. This model checkpoint also includes per tensor scales for FP8 quantized KV Cache, accessed through the kv cache dtype fp8 argument in vLLM. Usage and Creation Produced using AutoFP8 with calibration samples from ultrachat. Evaluation Open LLM Leaderboard evaluation scores Meta Llama 3 8B Instruct Meta Llama 3 8B Instruct FP8 Meta Llama 3 8B Instruct FP8 KV (this model) : : : : : : : : gsm8k 5 shot 75.44 74.37 74.98
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy