Qwen2.5 VL 32B Instruct FP8 Dynamic Model Overview Model Architecture: Qwen2.5 VL 32B Instruct Input: Vision Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: 5/3/2025 Version: 1.0 Model Developers: BC Card Quantized version of Qwen/Qwen2.5 VL 32B Instruct. Model Optimizations This model was obtained by quantizing the weights of Qwen/Qwen2.5 VL 32B Instruct to FP8 data type, ready for inference with vLLM = 0.5.2. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy