InternVideo2.5 is a video multimodal large language model (MLLM, built upoon InternVL2.5) enhanced with
long and rich context (LRC) modeling
. It significantly improves upon existing MLLMs by enhancing their ability to perceive fine-grained details and capture long-form temporal structures. We achieve this through dense vision task annotations using direct preference optimization (TPO) and compact spatiotemporal representations via adaptive hierarchical token compression (HiCo).
📈 Performance
Model
MVBench
LongVideoBench
VideoMME(w/o sub)
InternVideo2.5
75.7
60.6
65.1
🚀 How to use the model
First, you need to install
flash attention2
and some other modules. We provide a simple installation example below:
from transformers import AutoModel, AutoTokenizer
# model setting
model_path = 'OpenGVLab/InternVideo2_5_Chat_8B'
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModel.from_pretrained(model_path, trust_remote_code=True).half().cuda()
image_processor = model.get_vision_tower().image_processor
# evaluation setting
max_num_frames = 512
generation_config = dict(
do_sample=False,
temperature=0.0,
max_new_tokens=1024,
top_p=0.1,
num_beams=1
)
video_path = "your_video.mp4"# single-turn conversation
question1 = "Describe this video in detail."
output1, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question1, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)
print(output1)
# multi-turn conversation
question2 = "How many people appear in the video?"
output2, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question2, chat_history=chat_history, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)
print(output2)
✏️ Citation
@article{wang2025internvideo,
title={InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling},
author={Wang, Yi and Li, Xinhao and Yan, Ziang and He, Yinan and Yu, Jiashuo and Zeng, Xiangyu and Wang, Chenting and Ma, Changlian and Huang, Haian and Gao, Jianfei and Dou, Min and Chen, Kai and Wang, Wenhai and Qiao, Yu and Wang, Yali and Wang, Limin},
journal={arXiv preprint arXiv:2501.12386},
year={2025}
}
Runs of OpenGVLab InternVideo2_5_Chat_8B on huggingface.co
3.0K
Total runs
0
24-hour runs
44
3-day runs
34
7-day runs
37
30-day runs
More Information About InternVideo2_5_Chat_8B huggingface.co Model
InternVideo2_5_Chat_8B huggingface.co is an AI model on huggingface.co that provides InternVideo2_5_Chat_8B's model effect (), which can be used instantly with this OpenGVLab InternVideo2_5_Chat_8B model. huggingface.co supports a free trial of the InternVideo2_5_Chat_8B model, and also provides paid use of the InternVideo2_5_Chat_8B. Support call InternVideo2_5_Chat_8B model through api, including Node.js, Python, http.
InternVideo2_5_Chat_8B huggingface.co is an online trial and call api platform, which integrates InternVideo2_5_Chat_8B's modeling effects, including api services, and provides a free online trial of InternVideo2_5_Chat_8B, you can try InternVideo2_5_Chat_8B online for free by clicking the link below.
OpenGVLab InternVideo2_5_Chat_8B online free url in huggingface.co:
InternVideo2_5_Chat_8B is an open source model from GitHub that offers a free installation service, and any user can find InternVideo2_5_Chat_8B on GitHub to install. At the same time, huggingface.co provides the effect of InternVideo2_5_Chat_8B install, users can directly use InternVideo2_5_Chat_8B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
InternVideo2_5_Chat_8B install url in huggingface.co: