Model Overview Model Architecture: Kimi K2.5 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.1.0 Transformers: 4.57.6 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark (V0.11.2) Quantized layers: layers.0.mlp , experts , shared experts , self attn Weight quantization: OCP MXFP4, Static; self attn Perchannel, FP8E4M3, Static Activation quantization: OCP MXFP4, Dynamic; self attn Pertoken, FP8E4M3, Dynamic Calibration Dataset: Pile This model was built with Kimi K2.5 model by applying AMD Quark for MXFP4 quantization and PTPC FP8 quantization. Model Quantization The model was quantized from moonshotai/Kimi K2.5 using AMD Quark. The weights and activations are quantized to MXFP4, and self attn layers are quantized to PTPC FP8. Quantization scripts: Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Evaluation The model was evaluated on GSM8K benchmarks. Accuracy Benchmark Kimi K2.5 Kimi K2.5 MXFP4 AttnFP8(this model) Recovery GSM8K (flexible extract) 94.09 93.56 99.44% Reproduction The GSM8K results were obtained using the lm evaluation harness framework, based on the Docker image vl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy