Use on OpenRouter · Release blog post · Laguna XS 2.1 collection Laguna XS 2.1 Laguna XS 2.1 is a 33B total parameter Mixture of Experts model with 3B activated parameters per token designed for agentic coding and long horizon work on a local machine. This model is an upgraded version of our Laguna XS.2 model with a +5.4% jump on SWE bench Multilingual as well as stronger performance on terminal style tasks. [!NOTE] This repository contains official GGUF conversions from our standard format release built for llama.cpp (and compatible with vLLM and SGLang). To use with Ollama, pull directly with ollama pull laguna xs 2.1 . llama.cpp support is not yet upstreamed. See below. Highlights Mixed SWA and global attention layout : Laguna XS 2.1 uses sigmoid gating with per layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 40 total layers) KV cache in FP8 : KV cache quantized to FP8, reducing memory per token Native reasoning support : Interleaved thinking between tool calls with support for enabling and disabling thinking per request Local ready : At 33B total parameters and 3B activated, Laguna XS 2.1 is compact enough to…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy