Base model: poolside/Laguna XS 2.1 Laguna XS 2.1 , self quantized to GGUF by Atomic Chat. Built straight from Poolside's original weights with a per tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. Highlights 33.4B parameters : the weights this repo quantizes. Context length : 262,144 tokens (256K), as published by Poolside. 40 layers : Mixture of Experts, hybrid sliding window (512) and global attention. Full imatrix ladder : every quant is calibrated with an importance matrix. Mixed SWA and global attention layout : Laguna XS 2.1 uses sigmoid gating with per layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 40 total layers). KV cache in FP8 : KV cache quantized to FP8, reducing memory per token. Native reasoning support : Interleaved thinking between tool calls with support for enabling and disabling thinking per request. [!NOTE] These GGUFs are self quantized from the original weights , not a repack. The importance matrix keeps low bit quants closer to the full precision model. [!IMPORTANT] Always pass jinja so the Laguna XS 2.1 chat template is applied. Without it…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy