Model Overview kimi k2.5 eagle3 is an Eagle3 MTP draft model for accelerating inference of Kimi K2.5, trained with TorchSpec — an online speculative decoding training framework that runs FSDP training and inference concurrently. If you find this draft model useful, please give our project TorchSpec a 🌟 on GitHub. Training data is available at lightseekorg/kimi mtp dataset. Training Setup Cluster : 4 nodes × 8× H200 (32 GPUs total) Training : 2 nodes (16 GPUs), FSDP Inference : 2 nodes (16 GPUs), Engine (TP=8 per node) Duration : ~14 hours per phase Training ran in two phases, each 20k steps (~300k samples): Phase 1 : Regenerated open perfectblend dataset Phase 2 : Mixed dataset (English, VL, Chinese, function call, agent, creative writing) All training responses were regenerated by Kimi K2.5 via Engine to match the base model's exact token distribution. Training Curves The plots show loss, token acceptance accuracy, and simulated accept length during training. Both eval sets contain 256 samples drawn from each phase's own training corpus. Phase 1 (steps 0 → 20k): Phase 2 (steps 20k → 40k): Performance The primary metric is accept length — the average number of tokens accepted per…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy