HIPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs
This work is a companion to our earlier report
KAT-V1: Kwai-AutoThink Technical Report
, where we first introduced the
AutoThink paradigm
for controllable reasoning. While KAT-V1 outlined the overall framework of
SFT + RL
for adaptive reasoning, this paper provides the
detailed algorithmic design
of that training recipe.
Overview
We introduce
HiPO (Hybrid Policy Optimization for Dynamic Reasoning in LLMs)
, a novel RL framework designed to enable models to decide when to “think” (i.e., Think-on)and when to skip reasoning (i.e., Think-off), thereby striking a balance between correctness and efficiency.
HIPO has two main components:
Hybrid Data Pipeline
– Collects both think-on and think-off responses, categorizes queries by difficulty, and uses a strong model (e.g., DeepSeek-V3) to generate explanations that justify mode choices.
Hybrid Reward System
– Combines rewards for both modes, with bias adjustment to prevent overuse of long reasoning and mode-aware advantage functions to align decisions with performance gains.
Experimental Findings
Think-on Only (Overthinking).
Training only on Think-on data makes the model reason on all problems, causing inefficiency.
GRPO.
Improves accuracy by
+3.1%
, but increases token length on simple tasks.
HiPO-1.7B huggingface.co is an AI model on huggingface.co that provides HiPO-1.7B's model effect (), which can be used instantly with this Kwaipilot HiPO-1.7B model. huggingface.co supports a free trial of the HiPO-1.7B model, and also provides paid use of the HiPO-1.7B. Support call HiPO-1.7B model through api, including Node.js, Python, http.
HiPO-1.7B huggingface.co is an online trial and call api platform, which integrates HiPO-1.7B's modeling effects, including api services, and provides a free online trial of HiPO-1.7B, you can try HiPO-1.7B online for free by clicking the link below.
Kwaipilot HiPO-1.7B online free url in huggingface.co:
HiPO-1.7B is an open source model from GitHub that offers a free installation service, and any user can find HiPO-1.7B on GitHub to install. At the same time, huggingface.co provides the effect of HiPO-1.7B install, users can directly use HiPO-1.7B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.