Kimi K2.7 Code NVFP4 🚀 Available now — try this model at cogito.decart.ai and via OpenRouter . Description: NVFP4 quantized version of moonshotai/Kimi K2.7 Code, quantized with NVIDIA Model Optimizer. The routed expert linear layers are quantized to NVFP4 (4 bit float, block size 16) with an FP8 KV cache; attention (MLA), shared experts, the layer 0 dense MLP, lm head , and the vision tower / mm projector remain BF16 — the same precision split as nvidia/Kimi K2.6 NVFP4. Ready for inference with vLLM on NVIDIA Blackwell. This is a community reproduction produced by Decart; it is not affiliated with or endorsed by NVIDIA or Moonshot AI. Third Party Community Consideration This model is not owned or developed by NVIDIA or Decart's base model providers. It is built to a third party's requirements; see the non Decart Kimi K2.7 Code Model Card. License/Terms of Use: Use of this model is governed by the license of the base model, moonshotai/Kimi K2.7 Code (Modified MIT). Model Architecture: Architecture Type: Transformers Network Architecture: DeepSeek V3 (MLA attention, 384 routed experts + 1 shared expert, 61 layers), wrapped with a vision tower + mm projector ( KimiK25ForConditionalGe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy