Qwen3.5 9B fp8 19G 12G memory decrease speedup 30% vllm serve can run lightvl Developed by Myself is a lightweight Vision Language Model (VLM) quantization toolkit supporting FP8, INT8, FP8 Block. It integrates with vLLM for high throughput inference and supports Qwen3 VL, Qwen3.5, InternVL Chat, and Gemma 4 models. fast quant your model step by step: 1、 pip3 install lightvl 2、 lightvl YOUR HF MODEL PATH 3、 the output quant model path is YOUR HF MODEL PATH fp8 [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Qwen3.5 Highlights Qwen3.5 features the following enhancement: Unified Vision Language Foundation : Early…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy