Qwen3 Omni 30B A3B Instruct — NVFP4 W4A4 (full thinker, awq clip text calibration) ModelOpt NVFP4 W4A4 quantization of Qwen/Qwen3 Omni 30B A3B Instruct. The entire thinker text body — attention QKV/O + MoE experts — is quantized to NVFP4 with FP8 per tensor input scales. Embeddings, norms, the MoE router ( mlp.gate ), lm head , the audio encoder, the vision encoder, the talker, and code2wav stay in BF16. Total checkpoint size 27.6 GiB (vs 66 GiB BF16) Size reduction ~58% Hardware requirement NVIDIA Blackwell (sm 100+) for native FP4 GEMM Kernel path FlashInfer Cutlass NvFp4 Linear + FlashInfer TRT LLM NvFp4 MoE Exported safetensors NaN bytes 0 (calibrated with the ModelOpt side fix; see Mitigations below) Accuracy — Daily Omni (full split, n=1197) Benchmarked on NVIDIA B200 (HBM3e, 183 GiB) via the in tree run qwen omni acc benchmark.py harness. BF16 baseline run on the same hardware, same harness, same prompt set. CUDA graphs enabled (no enforce eager ). Model Daily Omni overall n correct Δ vs BF16 Qwen/Qwen3 Omni 30B A3B Instruct (BF16) 0.694 831 / 1197 — This checkpoint (W4A4 NVFP4) 0.675 808 / 1197 1.9 pp Accuracy — OmniBench (n=200, via evalscope) Model OmniBench mean acc Δ vs…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy