Qwen3 4B Instruct 2507 NVFP4 NVFP4 quantized version of Qwen/Qwen3 4B Instruct 2507 produced with llmcompressor. Notes Quantization scheme: NVFP4 (linear layers, lm head excluded) Calibration samples: 512 Max sequence length during calibration: 2048 Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy