LFM2.5 8B A1B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Liquid AI's original weights with a per-tensor importance matrix. Runs fully offline.
Highlights
Sparse MoE
: 8.3B total parameters, only 1.5B active per token.
LFM2 hybrid architecture
: 24 layers (18 double-gated LIV convolution blocks + 6 GQA attention), built on LFM2 with extended pre-training and reinforcement learning.
On-device assistant
: designed to chain tool calls and follow complex instructions, with day-one support for llama.cpp, MLX, vLLM and SGLang.
Reasoning model
: assistant turns include an explicit chain of thought before the final answer.
128K context
, 128,000 vocabulary, trained on a 38 trillion token budget.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
LFM2.5 8B A1B chat template
is applied. Without it the model can emit malformed turns.
Model Overview
Property
Value
Base model
LiquidAI/LFM2.5-8B-A1B
Total / active parameters
8.3B total, 1.5B active (MoE)
Layers
24 (18 LIV conv + 6 GQA)
Context length
128,000
Architecture
LFM2.5 hybrid (built on LFM2, extended pre-training + RL)
This repo
GGUF quants (imatrix)
Scores are Liquid AI's published results for the base
LiquidAI/LFM2.5-8B-A1B
. Quantization preserves the large majority of this;
Q4_K_M
and up sit within a point or two of full precision.
Choosing a quant
Quant
Size
Notes
Q2_K
3.2 GB
Smallest. Minimal RAM, clear quality drop.
IQ3_M
3.8 GB
Beats Q3 at similar size thanks to imatrix. Best low-RAM pick.
Q3_K_M
4.1 GB
Low quality but usable.
Q3_K_L
4.4 GB
A step above Q3_K_M.
IQ4_XS
4.6 GB
Excellent quality for size. Recommended low-bit.
Q4_K_S
4.9 GB
Compact Q4, fast.
Q4_K_M
5.2 GB
Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL
5.2 GB
Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_S
5.9 GB
Higher quality.
Q5_K_M
6.0 GB
Higher quality, low loss.
Q6_K
7.0 GB
Near lossless.
Q8_0
9.0 GB
Effectively lossless, reference quality.
Pick the largest file that fits your (V)RAM with room for context.
Q4_K_M
or
UD-Q4_K_XL
is the sweet spot for most setups;
Q6_K
or
Q8_0
for maximum fidelity.
Get started
Run LFM2.5 8B A1B locally with:
Atomic Chat
:
the easiest path. Open the app, search
AtomicChat/lfm25-8b-a1b-GGUF
, pick a quant, hit
Use this model
.
lfm25-8b-a1b-GGUF huggingface.co is an AI model on huggingface.co that provides lfm25-8b-a1b-GGUF's model effect (), which can be used instantly with this AtomicChat lfm25-8b-a1b-GGUF model. huggingface.co supports a free trial of the lfm25-8b-a1b-GGUF model, and also provides paid use of the lfm25-8b-a1b-GGUF. Support call lfm25-8b-a1b-GGUF model through api, including Node.js, Python, http.
lfm25-8b-a1b-GGUF huggingface.co is an online trial and call api platform, which integrates lfm25-8b-a1b-GGUF's modeling effects, including api services, and provides a free online trial of lfm25-8b-a1b-GGUF, you can try lfm25-8b-a1b-GGUF online for free by clicking the link below.
AtomicChat lfm25-8b-a1b-GGUF online free url in huggingface.co:
lfm25-8b-a1b-GGUF is an open source model from GitHub that offers a free installation service, and any user can find lfm25-8b-a1b-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of lfm25-8b-a1b-GGUF install, users can directly use lfm25-8b-a1b-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.