Qwen3 VL Embedding 8B — AWQ INT4 (W4A16) 4 bit AWQ quantization of Qwen/Qwen3 VL Embedding 8B , exported in the compressed tensors format for fast serving with vLLM (Marlin INT4 kernels on Ampere/Ada/Hopper and Jetson Orin sm 87 ). The model produces dense multimodal embeddings (text, image, video) in a shared 4096 dim space using last token pooling . This quant keeps the entire vision tower in BF16 and only quantizes the language decoder, which preserves embedding quality (see Evaluation). Base model Qwen/Qwen3 VL Embedding 8B (8B, qwen3 vl ) Method AWQ (activation aware), llm compressor Scheme W4A16 , 4 bit int weights / 16 bit activations, asymmetric Granularity group wise, group size = 128 , pack quantized Kept in BF16 full vision tower ( visual. , merger , patch embed ), lm head Format compressed tensors (vLLM native) Embedding dim 4096 (Matryoshka: usable down to 64) Pooling last token, include prompt: true Size on disk ~5.7 GB (vs ~16 GB BF16, ≈64% smaller) What is and isn't quantized Only the language model Linear layers are quantized to INT4. Quantizing the vision encoder of a VLM embedding model causes severe "cosine drift", so the following are explicitly excluded and re…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy