Qwen3.5 4B quantized.w4a16 Model Overview Model Architecture: Qwen/Qwen3.5 4B Input: Text / Image Output: Text Model Optimizations: Weight quantization: INT4 Activation quantization: None Model size: 5.2 GB (reduced from 8.8 GB in BF16) Release Date: 2026 04 16 Version: 1.0 Model Developers: RedHatAI This model is a quantized version of Qwen/Qwen3.5 4B. Evaluation results and reproduction steps are provided below. Model Optimizations This model was obtained by quantizing the weights of Qwen/Qwen3.5 4B to INT4 data type while keeping activations in original precision, ready for inference with vLLM. This optimization reduces the model weights from 8.8 GB to 5.2 GB on disk (~41% reduction). The reduction is less than the theoretical 75% because the vision encoder, token embeddings, and linear attention layers remain in BF16. Only the weights of the linear operators within transformer blocks are quantized using LLM Compressor. The vision encoder, token embeddings, and linear attention layers are not quantized. Deployment Use with vLLM 1. Initialize vLLM server: Multimodal (vision + text): Text only (lower memory): 2. Send requests to the server: Creation This model was created by apply…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy