We are pleased to announce the release of
Ovis2.5
, the successor to Ovis2, designed for native-resolution visual perception and enhanced multimodal reasoning.
It integrates a native-resolution vision transformer (NaViT) that processes images at their original, variable resolutions, eliminating the need for fixed-resolution tiling and preserving both fine details and global layout—crucial for visually dense content such as charts and diagrams.
To strengthen reasoning, Ovis2.5 is trained not only on linear chain-of-thought (CoT) but also on reflective reasoning, including self-checking and revision.
This advanced capability is available at inference as an optional
thinking mode
, enabling users to trade latency for higher accuracy on complex inputs.
Building on these advances,
Ovis2.5-9B
achieves an average score of 78.3 on the OpenCompass multimodal evaluation suite (SOTA among open-source MLLMs under 40B parameters), while the lightweight
Ovis2.5-2B
scores 73.9, continuing the “small model, big performance” philosophy for resource-constrained scenarios.
Key Features
Native-Resolution Perception
— NaViT vision encoder preserves fine details and global structure without lossy tiling.
Deep-Reasoning Capability
— Optional
thinking mode
for self-checking and revision beyond linear CoT.
Chart & Document OCR
— State-of-the-art at its scale for complex chart analysis, document understanding (including tables and forms), and OCR.
Broad Task Coverage
— Demonstrates leading performance on image reasoning, video understanding, and grounding benchmarks, showcasing strong general multimodal capability.
Quick Inference
Below is a simple example demonstrating how to run Ovis2.5 with a single image input.
To enable grounding, end your prompt with
Please provide the bounding box coordinates.
(for boxes) or
Please provide the point coordinates.
(for points). To target a specific object, wrap its description in
<ref>
tags, e.g.:
Find the <ref>red apple</ref> in the image. Please provide the bounding box coordinates.
Coordinates are normalized to
[0,1)
with the origin
(0,0)
at the top-left corner of the image.
Point:
<point>(x,y)</point>
Bounding box:
<box>(x1,y1),(x2,y2)</box>
where
(x1,y1)
is top-left,
(x2,y2)
is bottom-right.
Multiple results can be listed in square brackets:
[<box>(...)</box>,<box>(...)</box> ]
Example:
The image features a serene scene with <ref>three birds</ref>[
<box>(0.401,0.526),(0.430,0.557)</box>,
<box>(0.489,0.494),(0.516,0.526)</box>,
<box>(0.296,0.529),(0.324,0.576)</box>
] flying in formation against a clear blue sky.
We evaluate Ovis2.5 using
VLMEvalKit
, as employed in the OpenCompass multimodal and reasoning evaluation suite.
Citation
If you find Ovis useful, please consider citing the paper
@article{lu2024ovis,
title={Ovis: Structural Embedding Alignment for Multimodal Large Language Model},
author={Shiyin Lu and Yang Li and Qing-Guo Chen and Zhao Xu and Weihua Luo and Kaifu Zhang and Han-Jia Ye},
year={2024},
journal={arXiv:2405.20797}
}
We used compliance-checking algorithms during the training process, to ensure the compliance of the trained model to the best of our ability. Due to the complexity of the data and the diversity of language model usage scenarios, we cannot guarantee that the model is completely free of copyright issues or improper content. If you believe anything infringes on your rights or generates improper content, please contact us, and we will promptly address the matter.
Runs of AIDC-AI Ovis2.5-9B on huggingface.co
18.8K
Total runs
0
24-hour runs
-26
3-day runs
465
7-day runs
14.3K
30-day runs
More Information About Ovis2.5-9B huggingface.co Model
Ovis2.5-9B huggingface.co is an AI model on huggingface.co that provides Ovis2.5-9B's model effect (), which can be used instantly with this AIDC-AI Ovis2.5-9B model. huggingface.co supports a free trial of the Ovis2.5-9B model, and also provides paid use of the Ovis2.5-9B. Support call Ovis2.5-9B model through api, including Node.js, Python, http.
Ovis2.5-9B huggingface.co is an online trial and call api platform, which integrates Ovis2.5-9B's modeling effects, including api services, and provides a free online trial of Ovis2.5-9B, you can try Ovis2.5-9B online for free by clicking the link below.
AIDC-AI Ovis2.5-9B online free url in huggingface.co:
Ovis2.5-9B is an open source model from GitHub that offers a free installation service, and any user can find Ovis2.5-9B on GitHub to install. At the same time, huggingface.co provides the effect of Ovis2.5-9B install, users can directly use Ovis2.5-9B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.