vLLM Fixed FP8 Release Runtime Compatibility Update This repository has been rebuilt as a vLLM compatible FP8 checkpoint while preserving the Qwopus MTP module. vLLM 0.21.0 validated FP8 E4M3 block 128 MTP tensors retained 30/30 vLLM benchmark complete What changed? Qwopus3.6 27B v2 FP8 is a quantized version of the original 16 bit model. This FP8 quantized version retains the MTP layers, enabling speculative decoding acceleration in runtimes that support Qwen style MTP. This release is packaged and validated for vLLM FP8 inference while keeping the model card focused on runtime compatibility, format details, and tested environment information. FP8 / vLLM Compatibility Matrix Area Value Quantization format quant method=fp8 , fmt=e4m3 , dynamic activations, block wise FP8 weights with weight block size=[128,128] Language model layers 64 Qwen3.6 language layers Tensor index 1606 tensors, 407 FP8 scale tensors MTP module 22 mtp. tensors retained; 7 MTP linear weights are FP8 quantized with matching scale tensors vLLM engine vLLM 0.21.0, loaded as quantization=fp8 Validated stack Transformers 5.9.0, PyTorch 2.11.0+cu130, safetensors 0.7.0, CUDA 13.0, NVIDIA driver 580.126.09 Test GPU N…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy