Model Overview Description: The NVIDIA Llama 3.1 8B Instruct FP8 model is the quantized version of the Meta's Llama 3.1 8B Instruct model, which is an auto regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Llama 3.1 8B Instruct FP8 model is quantized with TensorRT Model Optimizer. This model is ready for commercial and non commercial use. Third Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third party’s requirements for this application and use case; see link to Non NVIDIA (Meta Llama 3.1 8B Instruct) Model Card. License/Terms of Use: nvidia open model license llama3.1 Model Architecture: Architecture Type: Transformers Network Architecture: Llama3.1 Input: Input Type(s): Text Input Format(s): String Input Parameters: Sequences Other Properties Related to Input: Context length up to 128K Output: Output Type(s): Text Output Format: String Output Parameters: Sequences Other Properties Related to Output: N/A Software Integration: Supported Runtime Engine(s): Tensor(RT) LLM vLLM Supported Hardware Microarchitecture Compatibility: NVID…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy