RedHatAI/Kimi K2.6 NVFP4 Model Overview Model Architecture: moonshotai/Kimi K2.6 ( KimiK25ForConditionalGeneration ) Input: Text, image, and video Output: Text Weight Quantization: NVFP4 (FP4 tensor group quantization) Activation Quantization: NVFP4 (FP4 tensor group quantization) Release Date: 2026 04 30 Model Developers: RedHatAI This model is a quantized variant of moonshotai/Kimi K2.6, exported in compressed tensors format for vLLM deployment and evaluated on instruction following, reasoning, function calling, and agentic coding workloads. Model Optimizations This checkpoint applies NVFP4 quantization to transformer linear layers with group wise FP4 weights and activations, using FP8 scale tensors. The format is optimized for efficient low precision serving while preserving strong benchmark quality on Kimi K2.6 evaluations. The model is exported in compressed tensors format and is intended for OpenAI compatible inference with vLLM. Creation This model was quantized with LLM Compressor and exported as compressed tensors. The script below is a representative reference script aligned with recipe.yaml and the published quantization configuration. Reference quantization script (NVFP…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy