We are excited to release the
CapRL 2.0 series
:
CapRL-Qwen3VL-2B
and
CapRL-Qwen3VL-4B
. These models feature fewer parameters while delivering even more powerful captioning performance.
Notably,
CapRL-Qwen3VL-4B significantly outperforms both CapRL-Qwen2.5VL-3B and Qwen2.5VL-72B in captioning tasks
, establishing itself as the top-performing model for captioning within the CapRL series.
This leap in efficiency is driven by our upgraded training recipe, which includes a more rigorous QA data filter and a significantly more diverse image dataset. We welcome everyone to try them out!
When selecting between the available CapRL models, it's essential to consider the trade-off between performance and computational cost.
This guide will help you choose the most suitable model for your specific needs:
We are excited to introduce
CapRL-3B
, a lightweight 3B image captioner that achieves perception capabilities comparable to Qwen2.5-VL-72B.
This is the first study of applying Reinforcement Learning with Verifiable Rewards for the
open-ended and subjective image captioning task. Unlike traditional Supervised Fine-Tuning, which
can lead to models memorizing a limited set of annotated captions, our method allows the model to
explore and generate a broader range of creative and general descriptions.
CapRL is a new training paradigm featuring a decoupled two-stage pipeline. The initial
stage uses LVLMs to generate rich and accurate captions. Subsequently, the second stage evaluates
caption quality by using a vision-only LLM to perform the QA task. We also created a specific QA
curation pipeline to ensure the quality of the questions and answers used for the second stage.
By employing the CapRL training framework, initializing with the Qwen2.5-VL-3B model, and using a carefully
filtered 75K QA dataset as the training set, we obtained a highly capable captioner,
CapRL-3B
.
Key Features
Remarkable visual understanding for Chart, Infographics and Document
:
CapRL-3B
achieves perception accuracy and visual information coverage comparable to Qwen2.5-VL-72B.
Well-organized output
: The outputs of CapRL-3B are relatively well-structured, making them clear and easy to understand.
Detailed description for natural images
: The outputs of
CapRL-3B
can perfectly cover all valid visual information while containing fewer hallucinations.
Usage
If you want to use
CapRL-3B
for captioning, you can directly follow the exact same inference approach as in
Qwen2.5-VL-series
.
We recommend using
vLLM
to speed up inference.
Start an OpenAI API Service
Run the command below to start an OpenAI-compatible API service:
CapRL-Qwen3VL-4B huggingface.co is an AI model on huggingface.co that provides CapRL-Qwen3VL-4B's model effect (), which can be used instantly with this internlm CapRL-Qwen3VL-4B model. huggingface.co supports a free trial of the CapRL-Qwen3VL-4B model, and also provides paid use of the CapRL-Qwen3VL-4B. Support call CapRL-Qwen3VL-4B model through api, including Node.js, Python, http.
CapRL-Qwen3VL-4B huggingface.co is an online trial and call api platform, which integrates CapRL-Qwen3VL-4B's modeling effects, including api services, and provides a free online trial of CapRL-Qwen3VL-4B, you can try CapRL-Qwen3VL-4B online for free by clicking the link below.
internlm CapRL-Qwen3VL-4B online free url in huggingface.co:
CapRL-Qwen3VL-4B is an open source model from GitHub that offers a free installation service, and any user can find CapRL-Qwen3VL-4B on GitHub to install. At the same time, huggingface.co provides the effect of CapRL-Qwen3VL-4B install, users can directly use CapRL-Qwen3VL-4B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.