Model Overview Model Architecture: Kimi K2.7 Code Input: Text, Image, Video Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.2.3 PyTorch: 2.10.0 Transformers: 5.12.1 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark (V0.12) Weight quantization: OCP MXFP4, Static; self attn Perchannel, FP8E4M3, Static Activation quantization: OCP MXFP4, Dynamic; self attn Pertoken, FP8E4M3, Dynamic Excluded from quantization: MoE gates, lm head , vision tower and multimodal projector This model was built with the Kimi K2.7 Code model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from moonshotai/Kimi K2.7 Code using AMD Quark. The MoE/Linear weights and activations are quantized to OCP MXFP4, while the attention projections use FP8 (E4M3). The vision tower and multimodal projector are kept at BF16. Quantization script: Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Note: this model has 64 KV heads, which is incompatible with the AITER MLA kernel (supports 16 or 128 only). Disable AITER MLA when serving on ROCm: Evaluation The model was evaluated on the GSM8K benchma…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy