Qwen2.5 Coder 32B Instruct FP8 Dynamic Model Overview Model Architecture: Qwen2.5 Coder 72B Instruct Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: 2/24/2025 Version: 1.0 Model Developers: BC Card Quantized version of Qwen/Qwen2.5 Coder 32B Instruct. Model Optimizations This model was obtained by quantizing the weights of Qwen/Qwen2.5 Coder 32B to FP8 data type, ready for inference with vLLM = 0.5.2. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details. Qwen2.5 Coder Introduction Qwen2.5 Coder is the latest series of Code Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5 Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5 Coder brings the following improvements upon CodeQwen1.5: Significantly improvements in code generation , code reasoning and code fixing . Base on the strong Qwen2.5, we scale up the training tokens into 5.5 trillion including so…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy