Qwen2.5 VL 3B Instruct FP8 Dynamic Model Overview Model Architecture: Qwen2.5 VL 3B Instruct Input: Vision Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: 2/24/2025 Version: 1.0 Model Developers: Neural Magic Quantized version of Qwen/Qwen2.5 VL 3B Instruct. Model Optimizations This model was obtained by quantizing the weights of Qwen/Qwen2.5 VL 3B Instruct to FP8 data type, ready for inference with vLLM = 0.5.2. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details. Creation This model was created with llm compressor by running the code snippet below as part a multimodal announcement blog. Model Creation Code Evaluation The model was evaluated using mistral evals for vision related tasks and using lm evaluation harness for select text based benchmarks. The evaluations were conducted using the following commands: Evaluation Commands Vision Tasks vqav2 docvqa mathvista mmmu chartqa Text based Tasks MMLU MGSM Accuracy Category Metric Qwen/Qwen2.5 VL 3B Instruct nm testing/Qwen2.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy