kimi k2.5 eagle3 mla Model Overview kimi k2.5 eagle3 mla is an Eagle3 MTP draft model with MLA(Multi Latent Attention) for accelerating inference of Kimi K2.5, trained with TorchSpec an online speculative decoding training framework that runs FSDP training and inference concurrently. If you find this draft model useful, please give our project TorchSpec a star on GitHub. Why an MLA (Multi Latent Attention) Draft Model Compared with an MHA draft model, the MLA variant is a better fit for Kimi K2.5 deployment: Uses less KV cache, which reduces serving memory pressure. Matches Kimi K2.5's MLA architecture, so it fits more naturally into the inference engine's KV cache handling under different serving scenarios such as PD Disaggregation. Training Setup Cluster: 4 nodes x 8x H200 (32 GPUs total) Training: 2 nodes (16 GPUs), FSDP Inference: 2 nodes (16 GPUs), Engine (TP=8 per node) Duration: ~14 hours per phase: Dataset: Regenerated open perfectblend dataset All training responses were regenerated by Kimi K2.5 via Engine to match the base model's exact token distribution. Performance The primary metric is accept length the average number of tokens accepted per speculation step with topk=…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy