Qwen3 VL Embedding 2B AWQ 4bit Quantized Model Overview This repository contains a 4 bit AWQ derivative of Qwen/Qwen3 VL Embedding 2B prepared for direct vLLM deployment through the compressed tensors backend. What Was Quantized Quantization method: llm compressor AWQ ( W4A16 ASYM ) Export format: compressed tensors Runtime backend: vLLM compressed tensors Weight format: 4 bit grouped asymmetric integer weights Group size: 128 Calibration pipeline: layer sequential Quantized modules: text side Linear layers in the Qwen3 VL decoder Left unquantized: all model.visual modules and lm head Calibration Data This checkpoint was built from the same 1000 sample mixed retrieval manifest as the FP16 and NVFP4 workflow, but the final AWQ pass used 876 text only samples and skipped 124 image bearing rows because the vision stack remained excluded from quantization. Calibration sources: Polish text retrieval: mteb/MSMARCO PL , mteb/NQ PL , mteb/FiQA PL Multilingual text retrieval: MIRACL hard negative slices for en , de , es , fr , ja Multimodal retrieval in the master manifest: vidore/colpali train set and lmms lab/flickr30k Hard negative augmentation: MIRACL derived negatives Local Benchmark S…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy