You can now run Laguna S 2.1 in Unsloth Studio or llama.cpp To run or train any LLM with a open source UI, install Unsloth Studio via: macOS, Linux, WSL: Windows: Running Unsloth's UD Q4 K XL with llama.cpp The GGUFs in this repo are Unsloth Dynamic 2.0 quants (imatrix calibrated). Laguna support is available starting from llama.cpp release b10087. Use this release or a newer one. Download the UD Q4 K XL shards (~40GB, split into 3 files): Serve it with llama server (pass the first shard; the remaining shards load automatically): Or run a one off generation with llama cli : Use on OpenRouter · Use on Vercel AI Gateway · Release blog post Laguna S 2.1 Laguna S 2.1 is a 118B total parameter Mixture of Experts model with 8B activated parameters per token, designed for agentic coding and long horizon work. It sits between Laguna XS 2.1 (33B A3B) and Laguna M.1 (225B A23B) in the Laguna series and shares the family recipe: a token choice router with softplus gating over 256 routed experts plus one shared expert, grouped query attention, and interleaved full/sliding window attention. Highlights Mixed SWA and global attention layout : 48 layers in a 1:3 global to SWA ratio (12 global atte…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy