RedHatAI/Kimi K2.6 FP8 BLOCK Model Overview Model Architecture: moonshotai/Kimi K2.6 ( KimiK25ForConditionalGeneration ) Input: Text, image, and video Output: Text Weight Quantization: FP8 (block wise scaling) Activation Quantization: FP8 (dynamic grouped scaling) Release Date: 2026 04 29 Model Developers: RedHatAI This model is a quantized variant of moonshotai/Kimi K2.6, exported in compressed tensors format for vLLM deployment and evaluated on instruction following, reasoning, function calling, and agentic coding workloads. Model Optimizations This checkpoint applies FP8 block quantization to transformer linear layers and FP8 dynamic quantization to activations. The resulting representation is optimized for high throughput serving while maintaining strong benchmark retention on Kimi K2.6 evaluation suites. The model is exported in compressed tensors format and is intended for OpenAI compatible inference with vLLM. Creation This model was quantized with LLM Compressor and exported as compressed tensors. The script below is a representative reference script aligned with the published quantization configuration. Reference quantization script (FP8 block) Deployment Use with vLLM Eva…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy