Laguna XS 2.1
, quantized to
GGUF
by
Atomic Chat
with an importance matrix. Built straight from poolside's original weights. Runs fully offline on your machine.
Highlights
33B total / 3B active
Mixture-of-Experts for agentic coding and long-horizon work on a local machine.
Mixed attention layout:
40 layers, 10 global + 30 sliding-window (3:1 ratio), sigmoid gating with per-layer rotary scales.
Native interleaved reasoning
, enable or disable per request.
Upgraded from Laguna XS.2
: +5.4% on SWE-bench Multilingual and stronger terminal-style performance.
Laguna is a new architecture. It runs in
Atomic Chat
1.1.135+
out of the box, or in a build of
llama.cpp with Laguna support
(
PR #25165
). Stock
llama.cpp
releases do not load it yet. Always pass
--jinja
so the chat template is applied.
Model Overview
Property
Value
Base model
poolside/Laguna-XS-2.1
Total parameters
33B (3B active per token)
Architecture
Laguna MoE, mixed sliding-window/global attention
Experts
256 + 1 shared
Layers
40 (10 global, 30 sliding-window)
Sliding window
512 tokens
Context length
262,144
Optimizer
Muon
This repo
imatrix GGUF quants for llama.cpp, built from the original weights.
Scores are poolside's published results for the full-precision base
poolside/Laguna-XS-2.1
. The GGUF quants run the same model locally; lower bit-widths trade a little accuracy for size and speed.
Choosing a quant
All rungs are quantized with an importance matrix (imatrix) calibrated on a general-purpose dataset.
Reasoning is native and on by default. For agentic coding, keep reasoning enabled and preserve prior thinking blocks across turns.
Best practices
Parameter
Value
temperature
1.0
top_k
20
top_p
1.0
poolside's benchmark settings.
How these were made
Download poolside's official
Laguna-XS-2.1-BF16.gguf
.
Build an importance matrix with
llama-imatrix
on a general calibration set.
Quantize each rung with
llama-quantize --imatrix
from the BF16 GGUF.
License
Released by poolside under the OpenMDW-1.1 license, which permits free use, modification and redistribution with attribution. GGUF conversion by Atomic Chat. This is an unofficial community quantization and is not endorsed by poolside; the original
LICENSE.md
and notices of origin are retained in this repo.
Runs of AtomicChat Laguna-XS-2.1-GGUF on huggingface.co
620
Total runs
0
24-hour runs
7
3-day runs
-47
7-day runs
-7.6K
30-day runs
More Information About Laguna-XS-2.1-GGUF huggingface.co Model
Laguna-XS-2.1-GGUF huggingface.co is an AI model on huggingface.co that provides Laguna-XS-2.1-GGUF's model effect (), which can be used instantly with this AtomicChat Laguna-XS-2.1-GGUF model. huggingface.co supports a free trial of the Laguna-XS-2.1-GGUF model, and also provides paid use of the Laguna-XS-2.1-GGUF. Support call Laguna-XS-2.1-GGUF model through api, including Node.js, Python, http.
Laguna-XS-2.1-GGUF huggingface.co is an online trial and call api platform, which integrates Laguna-XS-2.1-GGUF's modeling effects, including api services, and provides a free online trial of Laguna-XS-2.1-GGUF, you can try Laguna-XS-2.1-GGUF online for free by clicking the link below.
AtomicChat Laguna-XS-2.1-GGUF online free url in huggingface.co:
Laguna-XS-2.1-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Laguna-XS-2.1-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Laguna-XS-2.1-GGUF install, users can directly use Laguna-XS-2.1-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.