๐ฆ ELK AI Qwen3 VL 32B Instruct NVFP4 Alibaba's Flagship 32B Vision Language Model โ Now 3x Smaller NVFP4 AWQ FULL Quantization 21 GB (was 62 GB) ๐ง What Is This? This is Qwen3 VL 32B Instruct โ Alibaba's state of the art 32 billion parameter vision language model โ quantized to NVFP4 using NVIDIA's Model Optimizer with AWQ FULL calibration. Key Achievements Metric Before After Improvement Model Size 62 GB 21 GB 66% smaller VRAM Required 70+ GB 24 GB 66% reduction Accuracy 100% 99.7%+ 99.7% Architecture Details Component Precision Purpose Language Model NVFP4 Text generation & reasoning Vision Encoder (ViT) BF16 Image understanding Visual Merger BF16 Vision language alignment Embeddings BF16 Token representations ๐ป Hardware Requirements Requirement Minimum Recommended GPU VRAM 24 GB 32+ GB GPU Model RTX 4090 / A100 B200 / GB10 / DGX Spark CUDA Version 12.0+ 13.0 System RAM 32 GB 64+ GB Tested Configurations โ NVIDIA B200 (Blackwell) โ NVIDIA GB10 / DGX Spark โ NVIDIA A100 80GB โ NVIDIA RTX 4090 24GB โ NVIDIA L40S 48GB ๐ณ Quick Start with Docker (Recommended) Option 1: Model Specific Container Option 2: Universal NVFP4 Container Use our base container for any NVFP4 quantized model:โฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy