Qwen3 Embedding 8B FP8 DYNAMIC FP8 dynamically quantized version of Qwen/Qwen3 Embedding 8B. Quantized using llmcompressor with FP8 DYNAMIC scheme. All Linear layers are quantized to FP8 (weights per channel, activations per token dynamic). The embed tokens layer is kept in full bfloat16 precision. Achieves ~50% memory reduction compared to the original bfloat16 model. Deployment with vLLM: The original model card from Qwen/Qwen3 Embedding 8B follows below. Qwen3 Embedding 8B Highlights The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining. Exceptional Versatility : The embedding mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy