Model Overview Model Architecture: Kimi K2.6 Input: Text, Image, Video Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.2.0 PyTorch : 2.9.1 Transformers : 5.8.1 Operating System(s): Linux Inference Engine: SGLang/vLLM Model Optimizer: AMD Quark (v0.11.1) Quantized layers: experts , shared experts Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic This model was built with Kimi K2.6 model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from a BF16 decompressed version of moonshotai/Kimi K2.6 using AMD Quark. The original checkpoint uses native INT4 (compressed tensors) quantization; it was first decompressed to BF16 before applying MXFP4 quantization. The weights and activations are quantized to MXFP4. Quantization scripts: Deployment Use with vLLM/SGLang This model can be deployed efficiently using the vLLM and SGLang backends. Evaluation The model was evaluated on gsm8k benchmarks using the vllm framework. Accuracy Benchmark Kimi K2.6 Kimi K2.6 MXFP4 (this model) Recovery GSM8K (flexible extract) 93.93 93.25 99.3% Reproduction The GSM8K results were obtained using the vLLM framework,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy