AtomicChat / Laguna-XS-2.1-GGUF

huggingface.co
Total runs: 620
24-hour runs: 0
7-day runs: -47
30-day runs: -7.6K
Model's Last Updated: July 23 2026
text-generation

Introduction of Laguna-XS-2.1-GGUF

Model Details of Laguna-XS-2.1-GGUF


Laguna XS 2.1

Laguna XS 2.1 , quantized to GGUF by Atomic Chat with an importance matrix. Built straight from poolside's original weights. Runs fully offline on your machine.

Highlights
  • 33B total / 3B active Mixture-of-Experts for agentic coding and long-horizon work on a local machine.
  • Mixed attention layout: 40 layers, 10 global + 30 sliding-window (3:1 ratio), sigmoid gating with per-layer rotary scales.
  • 256 experts + 1 shared expert , sliding window of 512 tokens.
  • 262,144-token context.
  • Native interleaved reasoning , enable or disable per request.
  • Upgraded from Laguna XS.2 : +5.4% on SWE-bench Multilingual and stronger terminal-style performance.

Laguna is a new architecture. It runs in Atomic Chat 1.1.135+ out of the box, or in a build of llama.cpp with Laguna support ( PR #25165 ). Stock llama.cpp releases do not load it yet. Always pass --jinja so the chat template is applied.

Model Overview
Property Value
Base model poolside/Laguna-XS-2.1
Total parameters 33B (3B active per token)
Architecture Laguna MoE, mixed sliding-window/global attention
Experts 256 + 1 shared
Layers 40 (10 global, 30 sliding-window)
Sliding window 512 tokens
Context length 262,144
Optimizer Muon
This repo imatrix GGUF quants for llama.cpp, built from the original weights.
Laguna XS 2.1 benchmarks

Scores are poolside's published results for the full-precision base poolside/Laguna-XS-2.1 . The GGUF quants run the same model locally; lower bit-widths trade a little accuracy for size and speed.

Choosing a quant

All rungs are quantized with an importance matrix (imatrix) calibrated on a general-purpose dataset.

Quant Size Notes
Q3_K_M 15 GB smallest, usable
Q4_K_M 19 GB fast, low memory
Q5_K_M 23 GB balanced
Q6_K 26 GB recommended sweet spot
Q8_0 34 GB closest to the original

Q6_K is the best quality/size balance for most setups. Use Q3_K_M / Q4_K_M on tighter memory; Q8_0 when you want maximum fidelity.

Get started
  • Atomic Chat : open the app (1.1.135+), search AtomicChat/Laguna-XS-2.1-GGUF , pick a quant, hit Use this model .
  • llama.cpp (build with Laguna support):
    llama-cli -m Laguna-XS-2.1-Q6_K.gguf --jinja \
        -p "Write a Python retry wrapper with exponential backoff." -n 512
    
  • llama.cpp server:
    llama-server -m Laguna-XS-2.1-Q6_K.gguf --jinja -c 8192
    # OpenAI-compatible endpoint at http://localhost:8080/v1/chat/completions
    

Reasoning is native and on by default. For agentic coding, keep reasoning enabled and preserve prior thinking blocks across turns.

Best practices
Parameter Value
temperature 1.0
top_k 20
top_p 1.0

poolside's benchmark settings.

How these were made
  1. Download poolside's official Laguna-XS-2.1-BF16.gguf .
  2. Build an importance matrix with llama-imatrix on a general calibration set.
  3. Quantize each rung with llama-quantize --imatrix from the BF16 GGUF.
License

Released by poolside under the OpenMDW-1.1 license, which permits free use, modification and redistribution with attribution. GGUF conversion by Atomic Chat. This is an unofficial community quantization and is not endorsed by poolside; the original LICENSE.md and notices of origin are retained in this repo.

Runs of AtomicChat Laguna-XS-2.1-GGUF on huggingface.co

620
Total runs
0
24-hour runs
7
3-day runs
-47
7-day runs
-7.6K
30-day runs

More Information About Laguna-XS-2.1-GGUF huggingface.co Model

More Laguna-XS-2.1-GGUF license Visit here:

https://choosealicense.com/licenses/openmdw-1.1

Laguna-XS-2.1-GGUF huggingface.co

Laguna-XS-2.1-GGUF huggingface.co is an AI model on huggingface.co that provides Laguna-XS-2.1-GGUF's model effect (), which can be used instantly with this AtomicChat Laguna-XS-2.1-GGUF model. huggingface.co supports a free trial of the Laguna-XS-2.1-GGUF model, and also provides paid use of the Laguna-XS-2.1-GGUF. Support call Laguna-XS-2.1-GGUF model through api, including Node.js, Python, http.

Laguna-XS-2.1-GGUF huggingface.co Url

https://huggingface.co/AtomicChat/Laguna-XS-2.1-GGUF

AtomicChat Laguna-XS-2.1-GGUF online free

Laguna-XS-2.1-GGUF huggingface.co is an online trial and call api platform, which integrates Laguna-XS-2.1-GGUF's modeling effects, including api services, and provides a free online trial of Laguna-XS-2.1-GGUF, you can try Laguna-XS-2.1-GGUF online for free by clicking the link below.

AtomicChat Laguna-XS-2.1-GGUF online free url in huggingface.co:

https://huggingface.co/AtomicChat/Laguna-XS-2.1-GGUF

Laguna-XS-2.1-GGUF install

Laguna-XS-2.1-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Laguna-XS-2.1-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Laguna-XS-2.1-GGUF install, users can directly use Laguna-XS-2.1-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Laguna-XS-2.1-GGUF install url in huggingface.co:

https://huggingface.co/AtomicChat/Laguna-XS-2.1-GGUF

Url of Laguna-XS-2.1-GGUF

Laguna-XS-2.1-GGUF huggingface.co Url

Provider of Laguna-XS-2.1-GGUF huggingface.co

AtomicChat
ORGANIZATIONS

Other API from AtomicChat

huggingface.co

Total runs: 431
Run Growth: -206
Growth Rate: -52.15%
Updated:July 23 2026
huggingface.co

Total runs: 32
Run Growth: -42
Growth Rate: -135.48%
Updated:July 28 2026