Use on OpenRouter · Release blog post Laguna XS 2.1 Laguna XS 2.1 is a 33B total parameter Mixture of Experts model with 3B activated parameters per token designed for agentic coding and long horizon work on a local machine. This model is an upgraded version of our Laguna XS.2 model with a +5.4% jump on SWE bench Multilingual as well as stronger performance on terminal style tasks. [!NOTE] For more details on how we train, including on data automixing and async off policy agent RL, check out our recent technical report. Highlights Mixed SWA and global attention layout : Laguna XS 2.1 uses sigmoid gating with per layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 40 total layers) KV cache in FP8 : KV cache quantized to FP8, reducing memory per token Native reasoning support : Interleaved thinking between tool calls with support for enabling and disabling thinking per request Local ready : At 33B total parameters and 3B activated, Laguna XS 2.1 is compact enough to run on a Mac with 36 GB of RAM. Available on Ollama and llama.cpp. High quality FP8, NVFP4 and INT4 quantized variants available (see the collection) OpenM…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy