Qwen3 Omni 30B A3B Instruct NVFP4 (W4A8) Pre quantized NVFP4 version of Qwen/Qwen3 Omni 30B A3B Instruct for deployment on NVIDIA Blackwell GPUs (RTX 5090, B100, B200). Key Features 4 bit weight, 8 bit activation (W4A8) quantization using NVIDIA ModelOpt NVFP4 Pre quantized checkpoint — weights are packed as FP4 (uint8), not online quantization ~3.6x compression on MoE expert weights (54 GB → 15 GB) Same capabilities as the original: text, image, audio, video input → text + speech output Quantization Details Component Precision Quantized? Thinker MoE experts (gate up proj, down proj) NVFP4 packed uint8 Yes — pre quantized Thinker attention (q/k/v/o proj) NVFP4 (calibrated) Yes Thinker lm head BF16 No Thinker MoE router gates BF16 No Audio Encoder BF16 No Vision Encoder BF16 No Talker (MoE) BF16 No Code2Wav BF16 No Weight Format MoE expert weights are stored as: gate up proj : packed uint8 (2× FP4 values per byte) gate up proj scale : float8 e4m3fn (per block of 16 FP8 scale) gate up proj scale 2 : bfloat16 (global per tensor scale) Same pattern for down proj Quantization Config Memory Comparison Config Checkpoint Size Notes BF16 (original) ~60 GB Full precision FP8 (ModelOpt) ~40 G…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy