DeepSeek Coder V2 Lite Instruct FP8 Model Overview Model Architecture: DeepSeek Coder V2 Lite Instruct Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta Llama 3 7B Instruct, this models is intended for assistant like chat. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 7/18/2024 Version: 1.0 License(s): deepseek license Model Developers: Neural Magic Quantized version of DeepSeek Coder V2 Lite Instruct. It achieves an average score of 79.60 on the HumanEval+ benchmark, whereas the unquantized model achieves 79.33. Model Optimizations This model was obtained by quantizing the weights and activations of DeepSeek Coder V2 Lite Instruct to FP8 data type, ready for inference with vLLM = 0.5.2. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy