Valley is a cutting-edge multimodal large model designed to handle a variety of tasks involving text, images, and video data, which is developed by ByteDance. Our model
Achieved the best results in the inhouse e-commerce and short-video benchmarks, much better then other SOTA opensource models.
Demonstrated comparatively outstanding performance in the OpenCompass Benchmark.
Release
[2025/11/27] 🔥🔥🔥 We have released the technical report of Valley2.5! Check out the full paper here:
Valley2.5 Technical Report
.
[2025/10/26] 🔥🔥🔥 Update
Valley2.5
, significantly enhance multimodal understanding and reasoning capabilities, achieving 74.3 on OpenCompass Multi-modal Academic Leaderboard!
[2025/02/15] 🔥 Update
Valley2-DPO
, achieve 69.6 on OpenCompass Multi-modal Academic Leaderboard and update AutoModel usage for checkpoints.
For the LLM, we select Qwen3-8B-Base, chosen for its strong reasoning and language comprehension abilities. The Vision Encoder leverages Qwen2-VL-ViT, capable of processing dynamic-resolution inputs—a more robust alternative to the commonly used tiling approach when dealing with images of extreme aspect ratios. The Projector employs a 2×2 pixelshuffle downsampling on visual tokens, followed by a two-layer MLP with a 64k hidden dimension, providing high alignment capacity between modalities.
This architectural design ensures that Valley2.5 achieves a balanced trade-off between representational power, computational efficiency, and multimodal adaptability.
All of our open-source models are licensed under the
Apache-2.0
license.
We are Hiring
The Data-Ecommerce-Platform Governance-Basic Algorithms Team focuses on the research and development of multi-modal large model algorithms and foundational algorithms, continuously delving deeply into this field. Our mission is to optimize algorithms and collaborate with business teams to comprehensively govern the quality and ecosystem of ByteDance's e-commerce products. Currently, the team has a strong demand for foundational algorithm expertise in NLP, CV, and multimodal technologies. We welcome inquiries and look forward to working on challenging projects with talented individuals like you!
@article{wu2025valley2,
title={Valley2: Exploring Multimodal Models with Scalable Vision-Language Design},
author={Wu, Ziheng and Chen, Zhenghao and Luo, Ruipu and Zhang, Can and Gao, Yuan and He, Zhentao and Wang, Xian and Lin, Haoran and Qiu, Minghui},
journal={arXiv preprint arXiv:2501.05901},
year={2025}
}
Runs of bytedance-research Valley2.5 on huggingface.co
25
Total runs
0
24-hour runs
1
3-day runs
4
7-day runs
17
30-day runs
More Information About Valley2.5 huggingface.co Model
Valley2.5 huggingface.co is an AI model on huggingface.co that provides Valley2.5's model effect (), which can be used instantly with this bytedance-research Valley2.5 model. huggingface.co supports a free trial of the Valley2.5 model, and also provides paid use of the Valley2.5. Support call Valley2.5 model through api, including Node.js, Python, http.
Valley2.5 huggingface.co is an online trial and call api platform, which integrates Valley2.5's modeling effects, including api services, and provides a free online trial of Valley2.5, you can try Valley2.5 online for free by clicking the link below.
bytedance-research Valley2.5 online free url in huggingface.co:
Valley2.5 is an open source model from GitHub that offers a free installation service, and any user can find Valley2.5 on GitHub to install. At the same time, huggingface.co provides the effect of Valley2.5 install, users can directly use Valley2.5 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.