Qwen3 VL Embedding 8B W8A8 W8A8 INT8 quantization of Qwen/Qwen3 VL Embedding 8B — the 1 multimodal embedding model on MMEB V2 (77.9). Quantization Details Property Value Method GPTQ W8A8 INT8 (weights INT8 per channel, activations INT8 dynamic per token) Format compressed tensors (vLLM native) Tool llm compressor Calibration 512 samples from ultrachat 200k (train sft split), max 2048 tokens Vision encoder Kept in BF16 (not quantized) Non linear weights Preserved exactly (norms, biases, embed tokens) Model size ~10.5 GB (down from ~16 GB BF16) What's NOT quantized Vision encoder ( model.visual. ) — kept in full BF16 All RMSNorm weights Embedding table ( embed tokens ) lm head (not used for embeddings) No SmoothQuant was applied — it corrupts norm weights which destroys embedding quality. Serving with vLLM Tested with vLLM 0.17.1 on RTX 3090 (24GB). Uses ~10.4 GB VRAM for model + ~10.9 GB KV cache at 0.90 GPU utilization. API Usage Original Model Architecture : Qwen3 VL (8B params, 36 layers, 4096 dim embeddings) Context : 32K tokens Modalities : Text, images, screenshots, video Benchmarks : MMEB V2 77.9, MMTEB retrieval 81.08 MRL : Supports custom embedding dimensions (64 4096) See…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy