AxionML Qwen3.5 4B NVFP4 Developed by AxionML for open source serving and deployment use cases. Part of AxionML's effort to provide ready to serve quantized models for the community. This is an NVFP4 quantized version of Qwen/Qwen3.5 4B (4B parameters), quantized using NVIDIA TensorRT Model Optimizer. Weights and activations of linear layers are quantized to FP4, reducing disk size and GPU memory by ~4x compared to BF16. About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16 element micro blocks, so that 4 bit stored values remain numerically useful for neural network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power of two only E8M0) enables fractional scales and error minimizing scale selection strategies such as dual pass evaluation comparing "map max to 6" versus "map max to 4 with clipping." On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher precision FP3…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy