Laguna XS 2.1
, quantized to
MLX (5-bit)
by
Atomic Chat
for Apple Silicon. Built straight from poolside's original weights. Runs fully offline on your Mac.
Highlights
33B total / 3B active
Mixture-of-Experts for agentic coding and long-horizon work on a local machine.
Mixed attention layout:
40 layers, 10 global + 30 sliding-window (3:1 ratio), sigmoid gating with per-layer rotary scales.
Native interleaved reasoning
, enable or disable per request.
Upgraded from Laguna XS.2
: +5.4% on SWE-bench Multilingual and stronger terminal-style performance.
These are
MLX
builds for Apple Silicon (M-series), quantized from the original weights, not a repack. Laguna's architecture runs on
mlx-vlm
(0.6.3+) as a text model; stock
mlx-lm
does not yet include it.
Model Overview
Property
Value
Base model
poolside/Laguna-XS-2.1
Total parameters
33B (3B active per token)
Architecture
Laguna MoE, mixed sliding-window/global attention
Experts
256 + 1 shared
Layers
40 (10 global, 30 sliding-window)
Sliding window
512 tokens
Context length
262,144
Optimizer
Muon
This repo
MLX quants (3-8 bit) for Apple Silicon, built from the original weights with mlx-vlm.
Scores are poolside's published results for the full-precision base
poolside/Laguna-XS-2.1
. The MLX quants run the same model locally; lower bit-widths trade a little accuracy for size and speed.
This quant
This repo is the
5-bit
MLX build (~21 GB). The full ladder (5/6/8-bit) lives in the
Laguna XS 2.1 collection
.
Get started
Atomic Chat
:
open the app, search
AtomicChat/Laguna-XS-2.1-MLX-5bit
, pick a quant, hit
Use this model
.
python -m mlx_vlm server --model AtomicChat/Laguna-XS-2.1-MLX-5bit-5bit --host 0.0.0.0 --port 8080
# POST http://localhost:8080/v1/chat/completions with "model": "6bit"
Reasoning is native and on by default. Start the server with
--enable-thinking
(optionally
--thinking-budget N
) to keep it; omit the flag for direct, non-reasoning replies.
Best practices
Parameter
Value
temperature
1.0
top_k
20
top_p
1.0
poolside's benchmark settings. For agentic coding, keep reasoning enabled and preserve prior thinking blocks across turns.
Quantize each rung with
python -m mlx_vlm convert --hf-path poolside/Laguna-XS-2.1 -q --q-bits <N> --q-group-size 64
.
License
Released by poolside under the OpenMDW-1.1 license, which permits free use, modification and redistribution with attribution. MLX conversion by Atomic Chat. This is an unofficial community quantization and is not endorsed by poolside; the original
LICENSE.md
and notices of origin are retained in each quant folder.
Runs of AtomicChat Laguna-XS-2.1-MLX-5bit on huggingface.co
381
Total runs
-2
24-hour runs
-9
3-day runs
-68
7-day runs
-1.4K
30-day runs
More Information About Laguna-XS-2.1-MLX-5bit huggingface.co Model
Laguna-XS-2.1-MLX-5bit huggingface.co is an AI model on huggingface.co that provides Laguna-XS-2.1-MLX-5bit's model effect (), which can be used instantly with this AtomicChat Laguna-XS-2.1-MLX-5bit model. huggingface.co supports a free trial of the Laguna-XS-2.1-MLX-5bit model, and also provides paid use of the Laguna-XS-2.1-MLX-5bit. Support call Laguna-XS-2.1-MLX-5bit model through api, including Node.js, Python, http.
Laguna-XS-2.1-MLX-5bit huggingface.co is an online trial and call api platform, which integrates Laguna-XS-2.1-MLX-5bit's modeling effects, including api services, and provides a free online trial of Laguna-XS-2.1-MLX-5bit, you can try Laguna-XS-2.1-MLX-5bit online for free by clicking the link below.
AtomicChat Laguna-XS-2.1-MLX-5bit online free url in huggingface.co:
Laguna-XS-2.1-MLX-5bit is an open source model from GitHub that offers a free installation service, and any user can find Laguna-XS-2.1-MLX-5bit on GitHub to install. At the same time, huggingface.co provides the effect of Laguna-XS-2.1-MLX-5bit install, users can directly use Laguna-XS-2.1-MLX-5bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Laguna-XS-2.1-MLX-5bit install url in huggingface.co: