Laguna S 2.1
, self-quantized to NVFP4 by
Atomic Chat
. Built straight from Poolside's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
117.6B parameters
: the weights this repo quantizes.
Context length
: 1,048,576 tokens (1M), as published by Poolside.
48 layers
: Mixture-of-Experts, hybrid sliding-window (512) and global attention.
Full imatrix ladder
: every quant is calibrated with an importance matrix.
Mixed SWA and global attention layout
: 48 layers in a 1:3 global-to-SWA ratio (12 global attention layers, 36 sliding-window layers, window 512), with softplus attention gating and per-layer-type rotary scales.
Native reasoning support
: interleaved thinking between tool calls, with per-request control via enable_thinking.
Speculative decoding
: a trained DFlash draft model is available for lower-latency serving.
These NVFP4s are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Model Overview
Property
Value
Base model
poolside/Laguna-S-2.1
Parameters
117.6B
Layers
48
Experts
256 routed (top-10)
Sliding window
512 tokens
Context length
1,048,576 tokens (1M)
Vocabulary
100,352
Modalities
Text
Architecture
Mixture-of-Experts, 256 experts (top-10), hybrid sliding-window (512) and global attention, 48 attention heads over 8 KV heads,
LagunaForCausalLM
This repo
NVFP4 weights
Benchmarks
Benchmark
Score
Laguna S 2.1
70.2%
Tencent Hy3
71.7%
Inkling
63.8%
Nemotron 3 Ultra
56.4%
DeepSeek-V4-Pro Max
64.0%
Kimi K3
88.3%
Qwen 3.7 Max
74.5%
Muse Spark 1.1
80%
Claude Fable 5
88%
Scores are Poolside's published results for the base
poolside/Laguna-S-2.1
, not our own measurements. Quantization preserves the large majority of this;
Q4_K_M
and up stay close to full precision.
Get started
Atomic Chat
:
search
AtomicChat/Laguna-S-2.1-NVFP4
and hit
Use this model
.
Laguna-S-2.1-NVFP4 huggingface.co is an AI model on huggingface.co that provides Laguna-S-2.1-NVFP4's model effect (), which can be used instantly with this AtomicChat Laguna-S-2.1-NVFP4 model. huggingface.co supports a free trial of the Laguna-S-2.1-NVFP4 model, and also provides paid use of the Laguna-S-2.1-NVFP4. Support call Laguna-S-2.1-NVFP4 model through api, including Node.js, Python, http.
Laguna-S-2.1-NVFP4 huggingface.co is an online trial and call api platform, which integrates Laguna-S-2.1-NVFP4's modeling effects, including api services, and provides a free online trial of Laguna-S-2.1-NVFP4, you can try Laguna-S-2.1-NVFP4 online for free by clicking the link below.
AtomicChat Laguna-S-2.1-NVFP4 online free url in huggingface.co:
Laguna-S-2.1-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find Laguna-S-2.1-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of Laguna-S-2.1-NVFP4 install, users can directly use Laguna-S-2.1-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.