LiquidAI / LFM2.5-8B-A1B-Base

huggingface.co
Total runs: 1.8K
24-hour runs: 0
7-day runs: -501
30-day runs: -2.2K
Model's Last Updated: May 29 2026
text-generation

Introduction of LFM2.5-8B-A1B-Base

Model Details of LFM2.5-8B-A1B-Base

LFM2.5-8B-A1B-Base

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

  • On-device personal assistant : Designed to power real-life applications, chaining tool calls, and following complex instructions on all devices.
  • Compressed performance : Competitive with much larger dense and MoE models on instruction following and agentic tasks.
  • Unmatched throughput : Fastest in its size class on both CPU and GPU inference, with day-one support for llama.cpp, MLX, vLLM, and SGLang.

Find more information about LFM2.5-8B-A1B in our blog post .

image

* AA-Omniscience Index (higher is better) rewards correct answers and penalizes hallucinations. Scores range from -100 to 100. See more results on Artificial Analysis .

🗒️ Model Details
Model Parameters Description
LFM2.5-8B-A1B-Base 8.3B total / 1.5B active Pre-trained base model for fine-tuning
LFM2.5-8B-A1B 8.3B total / 1.5B active Reasoning-tuned general-purpose model

LFM2.5-8B-A1B is a general-purpose text-only model with the following features:

  • Total parameters : 8.3B
  • Active parameters : 1.5B
  • Number of layers : 24 (18 double-gated LIV conv + 6 GQA)
  • Training budget : 38 trillion tokens
  • Context length : 131,072
  • Vocabulary size : 128,000
  • Languages : English, Arabic, Chinese, French, German, Japanese, Korean, Portuguese, Spanish
  • Generation parameters : We recommend the following parameters:
    • temperature: 0.2
    • top_p: 80
    • repetition_penalty: 1.05
Model Description
LFM2.5-8B-A1B Original model checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM, and SGLang.
LFM2.5-8B-A1B-GGUF Quantized format for llama.cpp and compatible tools. Optimized for edge inference and local deployment.
LFM2.5-8B-A1B-ONNX ONNX Runtime format for cross-platform deployment.
LFM2.5-8B-A1B-MLX MLX format for Apple Silicon. Optimized for fast inference on Mac devices.

We recommend using LFM2.5-8B-A1B for agentic workflows, tool use, structured outputs, multilingual assistants, and on-device personal-assistant applications. It is not the best fit for heavy programming or knowledge-intensive question answering without retrieval.

🏃 Inference

LFM2.5-8B-A1B is supported by many inference frameworks. See the Inference documentation for the full list.

Name Description Docs Notebook
Transformers Simple inference with direct access to model internals. Link Colab link
vLLM High-throughput production deployments with GPU. Link Colab link
llama.cpp Cross-platform inference with CPU offloading. Link Colab link
MLX Apple's machine learning framework optimized for Apple Silicon. Link
LM Studio Desktop application for running LLMs locally. Link

Quick start with Transformers (compatible with transformers>=5.0.0 ):

from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_id = "LiquidAI/LFM2.5-8B-A1B-Base"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16",
#   attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

prompt = "What is C. elegans?"

input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    return_tensors="pt",
    tokenize=True,
).to(model.device)

output = model.generate(
    input_ids,
    do_sample=True,
    temperature=0.2,
    top_k=80,
    repetition_penalty=1.05,
    max_new_tokens=8192,
    streamer=streamer,
)
🔧 Fine-Tuning

We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.

Name Description Docs Notebook
CPT ( Unsloth ) Continued Pre-Training using Unsloth for text completion. Link Colab link
CPT ( Unsloth ) Continued Pre-Training using Unsloth for translation. Link Colab link
SFT ( Unsloth ) Supervised Fine-Tuning with LoRA using Unsloth. Link Colab link
SFT ( TRL ) Supervised Fine-Tuning with LoRA using TRL. Link Colab link
DPO ( TRL ) Direct Preference Optimization with LoRA using TRL. Link Colab link
GRPO ( Unsloth ) GRPO with LoRA using Unsloth. Link Colab link
GRPO ( TRL ) GRPO with LoRA using TRL. Link Colab link
📬 Contact
Citation
@article{liquidAI20268BA1B,
  author  = {Liquid AI},
  title   = {LFM2.5-8B-A1B: Personal Assistant On Your Laptop},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-8b-a1b},
}
@article{liquidai2025lfm2,
  title   = {LFM2 Technical Report},
  author  = {Liquid AI},
  journal = {arXiv preprint arXiv:2511.23404},
  year    = {2025}
}

Runs of LiquidAI LFM2.5-8B-A1B-Base on huggingface.co

1.8K
Total runs
0
24-hour runs
43
3-day runs
-501
7-day runs
-2.2K
30-day runs

More Information About LFM2.5-8B-A1B-Base huggingface.co Model

More LFM2.5-8B-A1B-Base license Visit here:

https://choosealicense.com/licenses/lfm1.0

LFM2.5-8B-A1B-Base huggingface.co

LFM2.5-8B-A1B-Base huggingface.co is an AI model on huggingface.co that provides LFM2.5-8B-A1B-Base's model effect (), which can be used instantly with this LiquidAI LFM2.5-8B-A1B-Base model. huggingface.co supports a free trial of the LFM2.5-8B-A1B-Base model, and also provides paid use of the LFM2.5-8B-A1B-Base. Support call LFM2.5-8B-A1B-Base model through api, including Node.js, Python, http.

LFM2.5-8B-A1B-Base huggingface.co Url

https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-Base

LiquidAI LFM2.5-8B-A1B-Base online free

LFM2.5-8B-A1B-Base huggingface.co is an online trial and call api platform, which integrates LFM2.5-8B-A1B-Base's modeling effects, including api services, and provides a free online trial of LFM2.5-8B-A1B-Base, you can try LFM2.5-8B-A1B-Base online for free by clicking the link below.

LiquidAI LFM2.5-8B-A1B-Base online free url in huggingface.co:

https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-Base

LFM2.5-8B-A1B-Base install

LFM2.5-8B-A1B-Base is an open source model from GitHub that offers a free installation service, and any user can find LFM2.5-8B-A1B-Base on GitHub to install. At the same time, huggingface.co provides the effect of LFM2.5-8B-A1B-Base install, users can directly use LFM2.5-8B-A1B-Base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

LFM2.5-8B-A1B-Base install url in huggingface.co:

https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-Base

Url of LFM2.5-8B-A1B-Base

LFM2.5-8B-A1B-Base huggingface.co Url

Provider of LFM2.5-8B-A1B-Base huggingface.co

LiquidAI
ORGANIZATIONS

Other API from LiquidAI

huggingface.co

Total runs: 119.0K
Run Growth: 76.2K
Growth Rate: 64.05%
Updated:April 02 2026
huggingface.co

Total runs: 104.2K
Run Growth: 78.6K
Growth Rate: 75.40%
Updated:March 30 2026
huggingface.co

Total runs: 100.4K
Run Growth: 58.9K
Growth Rate: 58.70%
Updated:February 13 2026
huggingface.co

Total runs: 94.9K
Run Growth: 25.8K
Growth Rate: 27.18%
Updated:March 30 2026
huggingface.co

Total runs: 20.2K
Run Growth: -12.8K
Growth Rate: -63.43%
Updated:March 30 2026
huggingface.co

Total runs: 15.0K
Run Growth: -24.0K
Growth Rate: -159.48%
Updated:March 30 2026
huggingface.co

Total runs: 11.9K
Run Growth: -170
Growth Rate: -1.43%
Updated:March 30 2026
huggingface.co

Total runs: 9.6K
Run Growth: -1.9K
Growth Rate: -19.52%
Updated:March 30 2026
huggingface.co

Total runs: 6.5K
Run Growth: 449
Growth Rate: 6.86%
Updated:March 30 2026
huggingface.co

Total runs: 3.8K
Run Growth: -692
Growth Rate: -18.35%
Updated:March 30 2026