dmis-lab / Qwen3-VL-8B-Thinking-MRPO

huggingface.co
Total runs: 17
24-hour runs: 0
7-day runs: 17
30-day runs: 17
Model's Last Updated: September 24 2026
image-text-to-text

Introduction of Qwen3-VL-8B-Thinking-MRPO

Model Details of Qwen3-VL-8B-Thinking-MRPO

Qwen3-VL-8B-Thinking-MRPO

MRPO is a novel reinforcement learning framework that improves medical multimodal reasoning by directly addressing failures in the reasoning process. It reshapes GRPO-style advantages using both answer-level and step-wise process rewards, assigning exponentially larger penalties to earlier invalid steps when the final answer is incorrect, thereby correcting early-stage failures before they cascade while preserving successful trajectories. By redistributing the learning signal according to where reasoning first fails, MRPO induces transferable reasoning that improves both reasoning quality and final answer accuracy across diverse medical VQA benchmarks.

Code: github

Project Page: page

Paper: Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Quick Start
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from PIL import Image
import torch

# Load the model (MRPO Qwen3-VL checkpoint; or a local trained checkpoint path)
model_path = "dmis-lab/Qwen3-VL-8B-Thinking-MRPO"
model = Qwen3VLForConditionalGeneration.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_path)

# Example usage (no system prompt; the chat template opens the reasoning with <think>)
image_path = "path/to/medical/image.jpg"
question = "What can you see in this medical image?"

question_text = (
    f"{question} Think step-by-step and enclose your reasoning in "
    "<think>...</think> tags. Then provide your answer in <answer>...</answer> tags."
)
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image_path},
            {"type": "text", "text": question_text},
        ],
    }
]

# Preparation for inference
text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(
    text=[text],
    images=[Image.open(image_path)],
    padding=True,
    padding_side="left",
    return_tensors="pt",
)
inputs = inputs.to(model.device)

# Inference (greedy decoding, matching inference.py)
generated_ids = model.generate(**inputs, use_cache=True, max_new_tokens=1024, do_sample=False)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)
Citation
@misc{jung2026breakingfailurecascadesstepaware,
      title={Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning}, 
      author={Junha Jung and Minbyul Jeong and Suhyeon Lim and Sungwook Jung and Jaehoon Yun and Taeyun Roh and Mujeen Sung and Jaewoo Kang},
      year={2026},
      eprint={2606.31825},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.31825}, 
}
License

This model is released under the Apache 2.0 license.

Runs of dmis-lab Qwen3-VL-8B-Thinking-MRPO on huggingface.co

17
Total runs
0
24-hour runs
3
3-day runs
17
7-day runs
17
30-day runs

More Information About Qwen3-VL-8B-Thinking-MRPO huggingface.co Model

More Qwen3-VL-8B-Thinking-MRPO license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-VL-8B-Thinking-MRPO huggingface.co

Qwen3-VL-8B-Thinking-MRPO huggingface.co is an AI model on huggingface.co that provides Qwen3-VL-8B-Thinking-MRPO's model effect (), which can be used instantly with this dmis-lab Qwen3-VL-8B-Thinking-MRPO model. huggingface.co supports a free trial of the Qwen3-VL-8B-Thinking-MRPO model, and also provides paid use of the Qwen3-VL-8B-Thinking-MRPO. Support call Qwen3-VL-8B-Thinking-MRPO model through api, including Node.js, Python, http.

Qwen3-VL-8B-Thinking-MRPO huggingface.co Url

https://huggingface.co/dmis-lab/Qwen3-VL-8B-Thinking-MRPO

dmis-lab Qwen3-VL-8B-Thinking-MRPO online free

Qwen3-VL-8B-Thinking-MRPO huggingface.co is an online trial and call api platform, which integrates Qwen3-VL-8B-Thinking-MRPO's modeling effects, including api services, and provides a free online trial of Qwen3-VL-8B-Thinking-MRPO, you can try Qwen3-VL-8B-Thinking-MRPO online for free by clicking the link below.

dmis-lab Qwen3-VL-8B-Thinking-MRPO online free url in huggingface.co:

https://huggingface.co/dmis-lab/Qwen3-VL-8B-Thinking-MRPO

Qwen3-VL-8B-Thinking-MRPO install

Qwen3-VL-8B-Thinking-MRPO is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-VL-8B-Thinking-MRPO on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-VL-8B-Thinking-MRPO install, users can directly use Qwen3-VL-8B-Thinking-MRPO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-VL-8B-Thinking-MRPO install url in huggingface.co:

https://huggingface.co/dmis-lab/Qwen3-VL-8B-Thinking-MRPO

Url of Qwen3-VL-8B-Thinking-MRPO

Qwen3-VL-8B-Thinking-MRPO huggingface.co Url

Provider of Qwen3-VL-8B-Thinking-MRPO huggingface.co

dmis-lab
ORGANIZATIONS

Other API from dmis-lab

huggingface.co

Total runs: 52.6K
Run Growth: -143.2K
Growth Rate: -272.17%
Updated:May 20 2021
huggingface.co

Total runs: 80
Run Growth: -1
Growth Rate: -1.25%
Updated:October 27 2021
huggingface.co

Total runs: 47
Run Growth: 33
Growth Rate: 70.21%
Updated:September 11 2024
huggingface.co

Total runs: 21
Run Growth: 8
Growth Rate: 38.10%
Updated:September 11 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 02 2025