DeepSeek V2.5 1210 FP8 Model Overview Model Architecture: DeepSeek V2.5 1210 Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: 3/1/2025 Version: 1.0 Model Developers: Neural Magic Quantized version of DeepSeek V2.5 1210. It achieves an average score of 77.8 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 77.82. Model Optimizations This model was obtained by quantizing the weights and activations to FP8 data type, ready for inference with vLLM = 0.5.2. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. The weights and activations of the linear operators within transformers blocks are quantized, except the MLP routers. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details. Creation This model was created with llm compressor by running the code snippet below with the following command: Evaluation The model was evaluated on OpenLLM Leaderboard V1 using t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy