Llama 3.1 8B Instruct FP8 Model creator: Meta Llama 3.1 Original model: Llama 3.1 8B Instruct Description This repo contains the Llama 3 8B Instruct model quantized to FP8 by FriendliAI, significantly enhancing its inference efficiency while maintaining high accuracy. Note that FP8 is only supported by NVIDIA Ada, Hopper, and Blackwell GPU architectures. Check out FriendliAI documentation for more details. License Refer to the license of the original model card. Compatibility This model is compatible with Friendli Container . Prerequisites Before you begin, make sure you have signed up for Friendli Suite. You can use Friendli Containers free of charge for four weeks. Prepare a Personal Access Token following this guide. Prepare a Friendli Container Secret following this guide. Preparing Personal Access Token PAT (Personal Access Token) is the user credential for for logging into our container registry. 1. Sign in Friendli Suite. 2. Go to User Settings Tokens and click 'Create new token' . 3. Save your created token value. Preparing Container Secret Container secret is a credential to launch our Friendli Container images. You should pass the container secret as an environment variab…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy