Qwen3 VL Embedding 8B FP8 This is an FP8 quantized version of Qwen/Qwen3 VL Embedding 8B, optimized for efficient inference with vLLM. Model Overview Attribute Value Base Model Qwen/Qwen3 VL Embedding 8B Quantization FP8 Dynamic (W8A8) Original Size ~16 GB (BF16) Quantized Size ~9 GB (FP8) Memory Savings ~45% Embedding Dimension 4096 Supported Inputs Text, Images, Videos, Multimodal Context Length 32K tokens Highlights Multimodal Versatility : Handles text, images, screenshots, and video inputs Efficient Inference : ~45% memory reduction with minimal accuracy loss vLLM Compatible : Works with vLLM's pooling runner for high throughput embedding No Calibration Required : Uses FP8 DYNAMIC scheme (data free quantization) Quantization Details Component Precision Notes Vision Encoder (ViT) BF16 Preserved for accuracy LLM Decoder Layers FP8 Quantized for efficiency Embeddings BF16 Preserved Scheme : FP8 DYNAMIC Weights: FP8 E4M3 (per channel quantization) Activations: Dynamic per token quantization at runtime Tool : llm compressor Calibration : None required (data free quantization) Hardware Requirements GPU : NVIDIA GPU with FP8 support (compute capability = 8.9) Blackwell: RTX 5090, RTX…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy