Youtu-VL
is a lightweight yet robust Vision-Language Model (VLM) built on the Youtu-LLM with 4B parameters. It pioneers Vision-Language Unified Autoregressive Supervision (VLUAS), which markedly strengthens visual perception and multimodal understanding. This enables a standard VLM to perform vision-centric tasks without task-specific additions. Across benchmarks, Youtu-VL stands out for its versatility, achieving competitive results on both vision-centric and general multimodal tasks.
✨ Key Features
Comprehensive Vision-Centric Capabilities
: The model demonstrates strong, broad proficiency across classic vision-centric tasks, delivering competitive performance in visual grounding, image classification, object detection, referring segmentation, semantic segmentation, depth estimation, object counting, and human pose estimation.
Promising Performance with High Efficiency
: Despite its compact 4B-parameter architecture, the model achieves competitive results across a wide range of general multimodal tasks, including general visual question answering (VQA), multimodal reasoning and mathematics, optical character recognition (OCR), multi-image and real-world understanding, hallucination evaluation, and GUI agent tasks.
Vision–Language Unified Autoregressive Supervision (VLUAS)
: Youtu-VL is built on the VLUAS paradigm to mitigate the text-dominant optimization bias in conventional VLMs, where visual signals are treated as passive conditions and fine-grained details are often dropped. Rather than using vision features only as inputs, Youtu-VL expands the text lexicon into a unified multimodal vocabulary through a learned visual codebook, turning visual signals into autoregressive supervision targets. Jointly reconstructing visual tokens and text explicitly preserves dense visual information while strengthening multimodal semantic understanding.
Vision-Centric Prediction with a Standard Architecture (no task-specific modules)
: Youtu-VL treats image and text tokens with equivalent autoregressive status, empowering it to perform vision-centric tasks for both dense vision prediction (e.g., segmentation, depth) and text-based prediction (e.g., grounding, detection) within a standard VLM architecture, eliminating the need for task-specific additions. This design yields a versitile general-purpose VLM, allowing a single model to flexibly accommodate a wide range of vision-centric and vsion-language requirements.
🏆 Model Performance
Vision-Centric Tasks
General Multimodal Tasks
🚀 Quickstart
Using Transformers to Chat
Ensure your Python environment has the
transformers
library installed and that the version meets the requirements.
Youtu-VL-4B-Instruct huggingface.co is an AI model on huggingface.co that provides Youtu-VL-4B-Instruct's model effect (), which can be used instantly with this tencent Youtu-VL-4B-Instruct model. huggingface.co supports a free trial of the Youtu-VL-4B-Instruct model, and also provides paid use of the Youtu-VL-4B-Instruct. Support call Youtu-VL-4B-Instruct model through api, including Node.js, Python, http.
Youtu-VL-4B-Instruct huggingface.co is an online trial and call api platform, which integrates Youtu-VL-4B-Instruct's modeling effects, including api services, and provides a free online trial of Youtu-VL-4B-Instruct, you can try Youtu-VL-4B-Instruct online for free by clicking the link below.
tencent Youtu-VL-4B-Instruct online free url in huggingface.co:
Youtu-VL-4B-Instruct is an open source model from GitHub that offers a free installation service, and any user can find Youtu-VL-4B-Instruct on GitHub to install. At the same time, huggingface.co provides the effect of Youtu-VL-4B-Instruct install, users can directly use Youtu-VL-4B-Instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Youtu-VL-4B-Instruct install url in huggingface.co: