TIGER-Lab / RationalRewards-8B-T2I

huggingface.co
Total runs: 153
24-hour runs: 0
7-day runs: 104
30-day runs: 110
Model's Last Updated: April 14 2026
image-to-text

Introduction of RationalRewards-8B-T2I

Model Details of RationalRewards-8B-T2I

TLDR: this is a reasoning reward model that supports text-to-image generation, from the following paper.


RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

Haozhe Wang 1 Cong Wei 2 Weiming Ren 2 Jiaming Liu 3 Fangzhen Lin 1 Wenhu Chen 2
1 HKUST 2 University of Waterloo 3 Alibaba

RationalRewards is a reasoning-based reward model and toolkit for visual generation. Instead of reducing preference into one opaque scalar, it generates explicit multi-dimensional critiques before scoring, turning reward models from passive evaluators into active optimization interfaces.

About the name: "Rational" means being reasonable, sensible, in Chinese, 理性的

RationalRewards supports optimization in complementary spaces :

  • train-time optimization through RL with structured, interpretable reward signals, and
  • test-time optimization through a Generate-Critique-Refine loop without parameter updates.
Key Results

Instantiated via PARROT on a Qwen3-VL-Instruct-8B backbone, RationalRewards achieves state-of-the-art preference prediction among open-source reward models and remains competitive with Gemini-2.5-Pro. As an RL reward, it consistently improves generators beyond scalar baselines across both text-to-image and image-editing tasks. Most interestingly, RationalRewards' test-time prompt tuning, requiring no parameter updates, matches or exceeds RL-based fine-tuning on several benchmarks.

RationalRewards teaser

Train-time RL and test-time prompt tuning with RationalRewards across visual generation benchmarks.

Why Reasoning Rewards?

Most reward models collapse instruction following, visual quality, composition, and plausibility into one scalar. This removes the structure of human judgment and often leads to brittle optimization. RationalRewards keeps those dimensions explicit so generators receive semantically grounded feedback about what to fix and why.

Why do reasoning rewards resist reward hacking?

Scalar rewards are vulnerable to reward hacking because they collapse rich judgment into one number that can rise even when outputs do not truly improve. RationalRewards introduces an implicit regularization: before giving scores, it must produce coherent, multi-dimensional critiques tied to concrete evaluation axes. This constrains optimization to evidence-backed reasoning and improves the monotonic relationship between reward and observed quality during RL.

Why are preference-trained rewards more stable than generic VLM judges?

Generic VLM judges can be strong analysts, but as reward functions they often show high-variance pointwise scoring across semantically similar samples. That variance becomes optimization noise in RL. PARROT trains RationalRewards directly for preference discrimination, yielding lower-variance, preference-aligned scores. The practical outcome is more stable optimization steps and better reward reliability, even with a smaller model footprint.

Why do reasoning rewards enable test-time scaling?

Reasoning feedback can be reused after generation, not only during training. In a Generate-Critique-Refine loop, RationalRewards critiques the produced image, identifies concrete deficiencies, and proposes targeted prompt updates. Unlike pre-hoc prompt enhancement that rewrites blindly, this is post-hoc and reactive to actual failures. That makes test-time compute more effective at eliciting latent generator capability, often approaching or surpassing RL fine-tuning gains without parameter updates.

RationalRewards usage overview

RationalRewards supports optimization in both parameter space (RL) and prompt space (test-time refinement).

Method: Preference-Anchored Rationalization (PARROT)

Human rationale annotation is expensive. PARROT recovers high-quality rationale supervision from preference-only data in three phases:

  1. Anchored generation: a teacher VLM proposes rationale candidates consistent with known labels.
  2. Consistency filtering: hallucinated or non-predictive rationales are removed.
  3. Distillation: a student model learns to critique-before-score without seeing labels.

This gives a practical path from abundant preference datasets to scalable reasoning supervision.

PARROT pipeline

PARROT pipeline: anchored rationale generation, consistency filtering, and distillation.

Empirical Evidence

RationalRewards strengthens both alignment quality and downstream optimization.

Preference prediction results

State-of-the-art preference prediction among open-source reward models.

Reward hacking analysis

Structured critique channels reduce shortcut exploitation compared with scalar-only rewards.

To better show optimization behavior, we also include diffusion RL training evolution results. The figure below visualizes how RationalRewards-guided training improves over time, illustrating that benefits are not only visible at the final checkpoint but emerge consistently throughout training. This constrast sharply with scalar rewards suffering reward hacking, as we demonstrate in Figure 12 in the paper.

Diffusion RL evolution

Evolution of diffusion RL performance under RationalRewards-guided optimization.

Test-time prompt tuning results

Generate-Critique-Refine at test time can match or exceed RL fine-tuning on several benchmarks.

Additional use cases

Additional qualitative use cases enabled by explicit reasoning feedback.

Citation
@article{rationalrewards2026,
  title   = {RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time},
  author  = {Haozhe Wang and Cong Wei and Weiming Ren and Jiaming Liu and Fangzhen Lin and Wenhu Chen},
  journal = {arXiv preprint},
  year    = {2026}
}

Runs of TIGER-Lab RationalRewards-8B-T2I on huggingface.co

153
Total runs
0
24-hour runs
0
3-day runs
104
7-day runs
110
30-day runs

More Information About RationalRewards-8B-T2I huggingface.co Model

More RationalRewards-8B-T2I license Visit here:

https://choosealicense.com/licenses/apache-2.0

RationalRewards-8B-T2I huggingface.co

RationalRewards-8B-T2I huggingface.co is an AI model on huggingface.co that provides RationalRewards-8B-T2I's model effect (), which can be used instantly with this TIGER-Lab RationalRewards-8B-T2I model. huggingface.co supports a free trial of the RationalRewards-8B-T2I model, and also provides paid use of the RationalRewards-8B-T2I. Support call RationalRewards-8B-T2I model through api, including Node.js, Python, http.

RationalRewards-8B-T2I huggingface.co Url

https://huggingface.co/TIGER-Lab/RationalRewards-8B-T2I

TIGER-Lab RationalRewards-8B-T2I online free

RationalRewards-8B-T2I huggingface.co is an online trial and call api platform, which integrates RationalRewards-8B-T2I's modeling effects, including api services, and provides a free online trial of RationalRewards-8B-T2I, you can try RationalRewards-8B-T2I online for free by clicking the link below.

TIGER-Lab RationalRewards-8B-T2I online free url in huggingface.co:

https://huggingface.co/TIGER-Lab/RationalRewards-8B-T2I

RationalRewards-8B-T2I install

RationalRewards-8B-T2I is an open source model from GitHub that offers a free installation service, and any user can find RationalRewards-8B-T2I on GitHub to install. At the same time, huggingface.co provides the effect of RationalRewards-8B-T2I install, users can directly use RationalRewards-8B-T2I installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

RationalRewards-8B-T2I install url in huggingface.co:

https://huggingface.co/TIGER-Lab/RationalRewards-8B-T2I

Url of RationalRewards-8B-T2I

RationalRewards-8B-T2I huggingface.co Url

Provider of RationalRewards-8B-T2I huggingface.co

TIGER-Lab
ORGANIZATIONS

Other API from TIGER-Lab

huggingface.co

Total runs: 2.9K
Run Growth: 2.5K
Growth Rate: 86.16%
Updated:October 14 2025
huggingface.co

Total runs: 925
Run Growth: 320
Growth Rate: 34.59%
Updated:December 06 2023
huggingface.co

Total runs: 877
Run Growth: 323
Growth Rate: 36.83%
Updated:December 06 2023
huggingface.co

Total runs: 520
Run Growth: 462
Growth Rate: 88.85%
Updated:January 09 2025
huggingface.co

Total runs: 42
Run Growth: -180
Growth Rate: -428.57%
Updated:July 15 2026
huggingface.co

Total runs: 40
Run Growth: -53
Growth Rate: -132.50%
Updated:July 15 2026
huggingface.co

Total runs: 30
Run Growth: -177
Growth Rate: -590.00%
Updated:July 15 2026
huggingface.co

Total runs: 22
Run Growth: -117
Growth Rate: -531.82%
Updated:July 15 2026
huggingface.co

Total runs: 19
Run Growth: -95
Growth Rate: -500.00%
Updated:July 15 2026
huggingface.co

Total runs: 19
Run Growth: -14
Growth Rate: -73.68%
Updated:November 08 2024