LFM2-VL-3B
is the newest and most capable model in
Liquid AI
's multimodal
LFM2-VL
series, designed to process text and images with variable resolutions.
Built on the
LFM2
backbone, it extends the architecture for higher-capacity reasoning and stronger visual understanding while retaining efficiency.
We are releasing the weights of the new
3B
checkpoint—offering higher performance across benchmarks while remaining optimized for scalable deployment.
Competitive multimodal performance
among lightweight open models.
Enhanced visual understanding and reasoning
, particularly on fine-grained perception tasks
Retains efficient inference
with the same flexible architecture and user-tunable speed-quality tradeoffs
Processes native resolutions up to 512×512
with intelligent patch-based handling for larger inputs
Due to their small size,
we recommend fine-tuning LFM2-VL models on narrow use cases
to maximize performance.
They were trained for instruction following and lightweight agentic flows.
Not intended for safety‑critical decisions.
Chat template
: LFM2-VL uses a ChatML-like chat template as follows:
<|startoftext|><|im_start|>system
You are a helpful multimodal assistant by Liquid AI.<|im_end|>
<|im_start|>user
<image>Describe this image.<|im_end|>
<|im_start|>assistant
This image shows a Caenorhabditis elegans (C. elegans) nematode.<|im_end|>
Images are referenced with a sentinel (
<image>
), which is automatically replaced with the image tokens by the processor.
You can apply it using the dedicated
.apply_chat_template()
function from Hugging Face transformers.
Architecture
Hybrid backbone
: Language model tower (LFM2-2.6B) paired with SigLIP2 NaFlex vision encoders (400M shape-optimized)
Native resolution processing
: Handles images up to 512×512 pixels without upscaling and preserves non-standard aspect ratios without distortion
Tiling strategy
: Splits large images into non-overlapping 512×512 patches and includes thumbnail encoding for global context
Here is an example of how to generate an answer with transformers in Python:
from transformers import AutoProcessor, AutoModelForImageTextToText
from transformers.image_utils import load_image
# Load model and processor
model_id = "LiquidAI/LFM2-VL-3B"
model = AutoModelForImageTextToText.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16"
)
processor = AutoProcessor.from_pretrained(model_id)
# Load image and create conversation
url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = load_image(url)
conversation = [
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": "What is in this image?"},
],
},
]
# Generate Answer
inputs = processor.apply_chat_template(
conversation,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
tokenize=True,
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
processor.batch_decode(outputs, skip_special_tokens=True)[0]
# This image captures a vibrant street scene in a Chinatown area. The focal point is a large red Chinese archway with gold and black accents, adorned with Chinese characters. Flanking the archway are two white stone lion statues, which are traditional guardians in Chinese culture.
You can directly run and test the model with this
Colab notebook
.
🔧 How to fine-tune
We recommend fine-tuning LFM2-VL models on your use cases to maximize performance.
Notebook
Description
Link
SFT (TRL)
Supervised Fine-Tuning (SFT) notebook with a LoRA adapter using TRL.
📈 Performance
Model
Average
MMStar
RealWorldQA
MM-IFEval
BLINK
MMBench (dev en)
OCRBench
POPE
InternVL3_5-2B
66.50
57.67
60.78
47.31
50.97
78.18
834.00
87.17
Qwen2.5-VL-3B
65.42
56.13
65.23
38.62
48.97
80.41
824.00
86.17
InternVL3-2B
67.44
61.10
65.10
38.49
53.10
81.10
831.00
90.10
SmolVLM2-2.2B
56.01
46.00
57.50
19.42
42.30
69.24
725.00
85.10
LFM2-VL-3B
69.00
57.73
71.37
51.83
51.03
79.81
822.00
89.01
More benchmark scores are reported in our
LFM2-VL-3B post
. We obtained the scores for competitive models using VLMEvalKit. Qwen3-VL-2B is not listed in the results table, as its release occurred the day before.
📬 Contact
If you are interested in custom solutions with edge deployment, please contact
our sales team
.
Runs of LiquidAI LFM2-VL-3B on huggingface.co
11.9K
Total runs
0
24-hour runs
0
3-day runs
946
7-day runs
-170
30-day runs
More Information About LFM2-VL-3B huggingface.co Model
LFM2-VL-3B huggingface.co is an AI model on huggingface.co that provides LFM2-VL-3B's model effect (), which can be used instantly with this LiquidAI LFM2-VL-3B model. huggingface.co supports a free trial of the LFM2-VL-3B model, and also provides paid use of the LFM2-VL-3B. Support call LFM2-VL-3B model through api, including Node.js, Python, http.
LFM2-VL-3B huggingface.co is an online trial and call api platform, which integrates LFM2-VL-3B's modeling effects, including api services, and provides a free online trial of LFM2-VL-3B, you can try LFM2-VL-3B online for free by clicking the link below.
LiquidAI LFM2-VL-3B online free url in huggingface.co:
LFM2-VL-3B is an open source model from GitHub that offers a free installation service, and any user can find LFM2-VL-3B on GitHub to install. At the same time, huggingface.co provides the effect of LFM2-VL-3B install, users can directly use LFM2-VL-3B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.