Use on OpenRouter · Release blog post Laguna XS 2.1 INT4 Laguna XS 2.1 INT4 is a 33B total parameter Mixture of Experts model with 3B activated parameters per token designed for agentic coding and long horizon work on a local machine. It uses Sliding Window Attention with per head gating in 30 out of 40 layers for fast inference and low KV cache requirements. [!NOTE] This is the INT4 variant with an FP8 quantized KV cache. The BF16, FP8 and NVFP4 variants are also available on Hugging Face. Highlights Mixed SWA and global attention layout : Laguna XS 2.1 uses sigmoid gating with per layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 40 total layers) KV cache in FP8 : KV cache quantized to FP8, reducing memory per token Native reasoning support : Interleaved thinking between tool calls with support for enabling and disabling thinking per request Local ready : At 33B total parameters and 3B activated, Laguna XS 2.1 is compact enough to run on a Mac with 36 GB of RAM. Available on Ollama and llama.cpp (BF16 and Q4\ K\ M only) OpenMDW 1.1 license : Use and modify the model and associated materials freely for commercial…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy