We release unofficial checkpoints for PPO, iterative DPO and rejection sampling (RAFT) trained from Qwen2.5-MATH-7B-base with rule-based RL, which are based on the success of Deepseek-R1-Zero and recent replications of PPO approach.
Evaluated on five widely-adopted benchmarks
AIME 2024
,
MATH 500
,
AMC
,
Minerva Math
,
OlympiadBench
, our
iterative DPO
and
RAFT
model achieve
significant enhancement compared to the base model and are comparable to the PPO approach.
Our models are trained by using the prompt set from the MATH training set and Numina Math.
Moreover, we provide a
detailed recipe
to reproduce the model. Enjoy!
Iterative DPO: Following the RLHF Workflow framework (
https://arxiv.org/pdf/2405.07863
), in each iteration, we sample multiple responses from the last trained policy, rank them via the ruled-based reward, and construct the preference pairs.
Then, we optimize the policy by minimizing the DPO loss and enter the next iteration.
Online iterative DPO can mitigate the issue of distribution shift and the limited coverage of offline data effectively
RLHFlow/Qwen2.5-7B-DPO is trained based on the fine-tuned model Qwen2.5-Math-7B-Base + SFT Warm-up.
More Information About Qwen2.5-7B-DPO huggingface.co Model
Qwen2.5-7B-DPO huggingface.co
Qwen2.5-7B-DPO huggingface.co is an AI model on huggingface.co that provides Qwen2.5-7B-DPO's model effect (), which can be used instantly with this RLHFlow Qwen2.5-7B-DPO model. huggingface.co supports a free trial of the Qwen2.5-7B-DPO model, and also provides paid use of the Qwen2.5-7B-DPO. Support call Qwen2.5-7B-DPO model through api, including Node.js, Python, http.
Qwen2.5-7B-DPO huggingface.co is an online trial and call api platform, which integrates Qwen2.5-7B-DPO's modeling effects, including api services, and provides a free online trial of Qwen2.5-7B-DPO, you can try Qwen2.5-7B-DPO online for free by clicking the link below.
RLHFlow Qwen2.5-7B-DPO online free url in huggingface.co:
Qwen2.5-7B-DPO is an open source model from GitHub that offers a free installation service, and any user can find Qwen2.5-7B-DPO on GitHub to install. At the same time, huggingface.co provides the effect of Qwen2.5-7B-DPO install, users can directly use Qwen2.5-7B-DPO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.