bytedance-research / Vidi1.5-9B

huggingface.co
Total runs: 42
24-hour runs: 2
7-day runs: 20
30-day runs: 20
Model's Last Updated: January 22 2026

Introduction of Vidi1.5-9B

Model Details of Vidi1.5-9B

Vidi: Large Multimodal Models for Video Understanding and Editing

Homepage: https://bytedance.github.io/vidi-website/

Github: https://github.com/bytedance/vidi

Demo: https://vidi.byteintl.com/

We introduce Vidi, a family of Large Multimodal Models (LMMs) for a wide range of video understanding and editing (VUE) scenarios. The first release focuses on temporal retrieval (TR), i.e., identifying the time ranges in input videos corresponding to a given text query.

This model is the Vidi1.5 model version for temporal retrieval.

Please find the inference, finetune and evaluation code on https://github.com/bytedance/vidi .

Citation

If you find Vidi useful for your research and applications, please cite using this BibTeX:

@article{Vidi2025vidi2,
          title={Vidi2: Large Multimodal Models for Video 
                  Understanding and Creation},
          author={Vidi Team, Celong Liu, Chia-Wen Kuo, Chuang Huang, 
                  Dawei Du, Fan Chen, Guang Chen, Haoji Zhang, 
                  Haojun Zhao, Lingxi Zhang, Lu Guo, Lusha Li, 
                  Longyin Wen, Qihang Fan, Qingyu Chen, Rachel Deng,
                  Sijie Zhu, Stuart Siew, Tong Jin, Weiyan Tao,
                  Wen Zhong, Xiaohui Shen, Xin Gu, Zhenfang Chen, Zuhua Lin},
          journal={arXiv preprint arXiv:2511.19529},
          year={2025}
}

@article{Vidi2025vidi,
          title={Vidi: Large Multimodal Models for Video 
                  Understanding and Editing},
          author={Vidi Team, Celong Liu, Chia-Wen Kuo, Dawei Du, 
                  Fan Chen, Guang Chen, Jiamin Yuan, Lingxi Zhang,
                  Lu Guo, Lusha Li, Longyin Wen, Qingyu Chen, 
                  Rachel Deng, Sijie Zhu, Stuart Siew, Tong Jin, 
                  Wei Lu, Wen Zhong, Xiaohui Shen, Xin Gu, Xing Mei, 
                  Xueqiong Qu, Zhenfang Chen},
          journal={arXiv preprint arXiv:2504.15681},
          year={2025}
}

Runs of bytedance-research Vidi1.5-9B on huggingface.co

42
Total runs
2
24-hour runs
5
3-day runs
20
7-day runs
20
30-day runs

More Information About Vidi1.5-9B huggingface.co Model

More Vidi1.5-9B license Visit here:

https://choosealicense.com/licenses/cc-by-nc-4.0

Vidi1.5-9B huggingface.co

Vidi1.5-9B huggingface.co is an AI model on huggingface.co that provides Vidi1.5-9B's model effect (), which can be used instantly with this bytedance-research Vidi1.5-9B model. huggingface.co supports a free trial of the Vidi1.5-9B model, and also provides paid use of the Vidi1.5-9B. Support call Vidi1.5-9B model through api, including Node.js, Python, http.

bytedance-research Vidi1.5-9B online free

Vidi1.5-9B huggingface.co is an online trial and call api platform, which integrates Vidi1.5-9B's modeling effects, including api services, and provides a free online trial of Vidi1.5-9B, you can try Vidi1.5-9B online for free by clicking the link below.

bytedance-research Vidi1.5-9B online free url in huggingface.co:

https://huggingface.co/bytedance-research/Vidi1.5-9B

Vidi1.5-9B install

Vidi1.5-9B is an open source model from GitHub that offers a free installation service, and any user can find Vidi1.5-9B on GitHub to install. At the same time, huggingface.co provides the effect of Vidi1.5-9B install, users can directly use Vidi1.5-9B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Vidi1.5-9B install url in huggingface.co:

https://huggingface.co/bytedance-research/Vidi1.5-9B

Url of Vidi1.5-9B

Provider of Vidi1.5-9B huggingface.co

bytedance-research
ORGANIZATIONS

Other API from bytedance-research