Fanar-Math-R1-GRPO
is a reasoning-optimized language model built on
QCRI/Fanar-1-9B-Instruct
. This version is fine-tuned using
Group Relative Policy Optimization (GRPO)
from the DeepSeekMath framework on the
AI-MO/NuminaMath-TIR
dataset. It is designed for step-by-step mathematical problem-solving with structured reasoning in both English and Arabic.
๐ Model Highlights
๐ Fine-tuned with
GRPO
, a sample-efficient reinforcement learning method
๐งฎ Specializes in
multi-step mathematical reasoning
๐ฌ Outputs responses in a structured conversational format using
<think>
and
<answer>
tags
๐ง Trained using
TRL
(
transformers
,
peft
, and
math_verify
)
๐ท๏ธ Useful for both instruction-following and math-heavy dialogue generation
@article{zhihong2024deepseekmath,
title={DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models},
author={Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Zhang, Mingchuan and Li, Y.K. and Wu, Y. and Guo, Daya},
journal={arXiv preprint arXiv:2402.03300},
year={2024}
}
TRL Library
@misc{vonwerra2022trl,
title={TRL: Transformer Reinforcement Learning},
author={von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouรฉdec, Quentin},
year={2022},
howpublished={\url{https://github.com/huggingface/trl}}
}
@misc{fanarllm2025,
title={Fanar: An Arabic-Centric Multimodal Generative AI Platform},
author={Fanar Team and Ummar Abbas and Mohammad Shahmeer Ahmad and Firoj Alam and Enes Altinisik and Ehsannedin Asgari and Yazan Boshmaf and Sabri Boughorbel and Sanjay Chawla and Shammur Chowdhury and Fahim Dalvi and Kareem Darwish and Nadir Durrani and Mohamed Elfeky and Ahmed Elmagarmid and Mohamed Eltabakh and Masoomali Fatehkia and Anastasios Fragkopoulos and Maram Hasanain and Majd Hawasly and Mus'ab Husaini and Soon-Gyo Jung and Ji Kim Lucas and Walid Magdy and Safa Messaoud and Abubakr Mohamed and Tasnim Mohiuddin and Basel Mousi and Hamdy Mubarak and Ahmad Musleh and Zan Naeem and Mourad Ouzzani and Dorde Popovic and Amin Sadeghi and Husrev Taha Sencar and Mohammed Shinoy and Omar Sinan and Yifan Zhang and Ahmed Ali and Yassine El Kheir and Xiaosong Ma and Chaoyi Ruan}},
year={2025},
url={https://arxiv.org/abs/2501.13944},
}
Fanar-Math-R1-GRPO huggingface.co is an AI model on huggingface.co that provides Fanar-Math-R1-GRPO's model effect (), which can be used instantly with this Omartificial-Intelligence-Space Fanar-Math-R1-GRPO model. huggingface.co supports a free trial of the Fanar-Math-R1-GRPO model, and also provides paid use of the Fanar-Math-R1-GRPO. Support call Fanar-Math-R1-GRPO model through api, including Node.js, Python, http.
Fanar-Math-R1-GRPO huggingface.co is an online trial and call api platform, which integrates Fanar-Math-R1-GRPO's modeling effects, including api services, and provides a free online trial of Fanar-Math-R1-GRPO, you can try Fanar-Math-R1-GRPO online for free by clicking the link below.
Omartificial-Intelligence-Space Fanar-Math-R1-GRPO online free url in huggingface.co:
Fanar-Math-R1-GRPO is an open source model from GitHub that offers a free installation service, and any user can find Fanar-Math-R1-GRPO on GitHub to install. At the same time, huggingface.co provides the effect of Fanar-Math-R1-GRPO install, users can directly use Fanar-Math-R1-GRPO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.