Use on OpenRouter · Use on Vercel AI Gateway · Release blog post Laguna S 2.1 NVFP4 Laguna S 2.1 NVFP4 is a 117.6B total parameter Mixture of Experts model with 8.5B activated parameters per token designed for agentic coding and long horizon work on a local machine. It uses Sliding Window Attention with per head gating in 36 out of 48 layers for fast inference and low KV cache requirements. Highlights Mixed SWA and global attention layout : Laguna S 2.1 uses softplus gating with per layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 48 total layers) KV cache in FP8 : KV cache quantized to FP8, reducing memory per token Native reasoning support : Interleaved thinking between tool calls with support for enabling and disabling thinking per request Local ready : At 117.6B total parameters and 8.5B activated, the NVFP4 weights are roughly 71 GB. Available on Ollama and llama.cpp (BF16 and Q4\ K\ M only) OpenMDW 1.1 license : Use and modify the model and associated materials freely for commercial and non commercial purposes (learn more about OpenMDW) Model overview Training: pre training, post training and reinforcement l…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy