LFM2.5 8B A1B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Liquid AI's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
8.5B parameters
: the weights this repo quantizes.
Context length
: 128,000 tokens (125K), as published by Liquid AI.
24 layers
: Mixture-of-Experts.
Full imatrix ladder
: every quant is calibrated with an importance matrix.
On-device personal assistant
: Designed to power real-life applications, chaining tool calls, and following complex instructions on all devices.
Compressed performance
: Competitive with much larger dense and MoE models on instruction following and agentic tasks.
Unmatched throughput
: Fastest in its size class on both CPU and GPU inference, with day-one support for llama.cpp, MLX, vLLM, and SGLang.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
LFM2.5 8B A1B chat template
is applied. Without it the model can emit malformed turns.
Scores are Liquid AI's published results for the base
LiquidAI/LFM2.5-8B-A1B
, not our own measurements. Quantization preserves the large majority of this;
Q4_K_M
and up stay close to full precision.
Beats Q3 at a similar size thanks to imatrix. Best low-RAM pick.
Q3_K_M
4.1 GB
Low quality but usable.
Q3_K_L
4.4 GB
A step above Q3_K_M.
IQ4_XS
4.6 GB
Excellent quality for size. Recommended low-bit.
Q4_K_S
4.9 GB
Compact 4-bit, fast.
Q4_K_M
5.2 GB
Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL
5.2 GB
Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_S
5.9 GB
Higher quality, slightly more compact than Q5_K_M.
Q5_K_M
6.0 GB
Higher quality, low loss.
Q6_K
7.0 GB
Near lossless, noticeably lighter than Q8_0.
Q8_0
9.0 GB
Effectively lossless, reference quality.
Pick the largest file that fits your (V)RAM with room for context.
Q4_K_M
or
UD-Q4_K_XL
is the sweet spot for most setups;
Q6_K
or
Q8_0
for maximum fidelity.
Get started
Run LFM2.5 8B A1B locally with:
Atomic Chat
:
the easiest path. Open the app, search
AtomicChat/lfm25-8b-a1b-GGUF
, pick a quant, hit
Use this model
.
LFM2.5-8B-A1B-GGUF huggingface.co is an AI model on huggingface.co that provides LFM2.5-8B-A1B-GGUF's model effect (), which can be used instantly with this AtomicChat LFM2.5-8B-A1B-GGUF model. huggingface.co supports a free trial of the LFM2.5-8B-A1B-GGUF model, and also provides paid use of the LFM2.5-8B-A1B-GGUF. Support call LFM2.5-8B-A1B-GGUF model through api, including Node.js, Python, http.
LFM2.5-8B-A1B-GGUF huggingface.co is an online trial and call api platform, which integrates LFM2.5-8B-A1B-GGUF's modeling effects, including api services, and provides a free online trial of LFM2.5-8B-A1B-GGUF, you can try LFM2.5-8B-A1B-GGUF online for free by clicking the link below.
AtomicChat LFM2.5-8B-A1B-GGUF online free url in huggingface.co:
LFM2.5-8B-A1B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find LFM2.5-8B-A1B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of LFM2.5-8B-A1B-GGUF install, users can directly use LFM2.5-8B-A1B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.