Fanar-1-9B
is a powerful Arabic-English LLM developed by
Qatar Computing Research Institute (QCRI)
at
Hamad Bin Khalifa University (HBKU)
, a member of Qatar Foundation for Education, Science, and Community Development. We continually pretrain the
google/gemma-2-9b
model on 1T Arabic and English tokens. We pay particular attention to the richness of the Arabic language by supporting Modern Standard Arabic (MSA) and a diverse set of Arabic dialects, including Gulf, Levantine, and Egyptian. Fanar models, through meticulous curation of the pretraining and instruction-tuning data, are aligned with Islamic values and Arab cultures.
The
instruction-tuned version
of
Fanar-1-9B
is a core component of the
Fanar GenAI platform
that offers a suite of capabilities including image generation, video and image understanding, deep thinking, advanced text-to-speech (TTS) and automatic-speech-recognition (ASR), attribution and fact-checking, Islamic RAG, among several other features.
We have published a comprehensive
report
with all the details regarding our Fanar GenAI platform. We also provide an API to our models and the GenAI platform (request access
here
).
Fanar-1-9B was continually pretrained on 1T tokens, with a balanced focus on Arabic and English: ~515B English tokens from a carefully curated subset of the
Dolma
dataset, 410B Arabic tokens that we collected, parsed, and filtered from a variety of sources, and 102B code tokens curated from
The Stack
dataset. Our codebase used the
LitGPT
framework.
Getting Started
Fanar-1-9B is compatible with the Hugging Face
transformers
library (≥ v4.40.0). Here's how to load and use the model:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "QCRI/Fanar-1-9B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
# prompt may be in Arabic or English
prompt = "ما هي عاصمة قطر؟"
inputs = tokenizer(prompt, return_tensors="pt", return_token_type_ids=False)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Intended Use
Fanar-1-9B is a base model and can be finetuned for a varierty of usecases such as:
Conversational agents (Arabic only or bilingual)
Cultural and dialectal question answering in Arabic
Educational, governmental, and civic NLP applications focused on the Arab world or Arabic-speaking audiences
Research on Arabic natural language generation and understanding
A finetuned version of Fanar-1-9B can be deployed as part of a broader AI system. Developers are encouraged to implement proper safeguards to ensure culturally respectful, accurate, and safe deployment. It should not be used to generate or spread
harmful, illegal, or misleading content.
Ethical Considerations & Limitations
Fanar-1-9B- is capable of generating fluent and contextually appropriate responses. However, as with any generative model there are uncertainities. The model may produce
biased, offensive, or incorrect outputs
. The model is
not suitable for high-stakes decision-making
(e.g., legal, medical, or financial advice). Though we have extensively tested Fanar-1-9B and attempted to mitigate these issues, we cannot redress every possible scenario. Thus, we advise developers to implement safety checks and perform domain-specific fine-tuning for sensitive use cases. Kindly refer to our
Terms of Service
and
Privacy Policy
.
The output generated by this model is not considered a statement of QCRI, HBKU, Qatar Foundation, MCIT or any other organization or individual.
Evaluation
Evaluation was conducted using a modified version of the LM Evaluation Harness and internal cultural alignment benchmarks.
Model
MMLU (5-shot)
MMMLU (Arabic) (0-shot)
ArabicMMLU (3-shot)
HellaSwag (0-shot)
PIQA (0-shot)
ARC Challenge (0-shot)
Belebele (Arabic) (3-shot)
ACVA (5-shot)
GSM8k
OALL (0-shot)
OALL v2 (0-shot)
Almieyar Arabic (3-shot)
Arab Cultural MCQ (3-shot)
AraDiCE PIQA (MSA) (0-shot)
AraDiCE PIQA(Egy) (0-shot)
AraDiCE PIQA(Lev) (0-shot)
AraDiCE ArabicMMLU(Egy) (0-shot)
AraDiCE ArabicMMLU(Lev) (0-shot)
Fanar-1-9B
71.33%
57.38%
67.42%
80.76%
81.66%
59.73%
79.31%
81.31%
45.79%
54.94%
63.20%
77.18%
72.30%
66.00%
62.19%
57.67%
55.79%
55.63%
AceGPT-v2-8B
63.55%
41.71%
58.55%
76.97%
80.03%
49.40%
60.61%
78.36%
10.92%
43.58%
47.00%
66.83%
67.50%
63.17%
61.48%
56.75%
43.40%
40.96%
gemma-2-9b
70.60%
54.04%
64.32%
79.82%
82.97%
65.53%
75.31%
79.66%
21.61%
50.24%
57.23%
73.82%
68.60%
63.98%
60.17%
58.05%
49.61%
47.15%
jais-adapted-13b
50.42%
34.01%
51.96%
78.02%
78.94%
48.55%
43.02%
73.52%
5.76%
40.79%
40.06%
62.34%
60.90%
65.02%
62.19%
59.25%
38.24%
37.93%
jais-family-6p7b
32.50%
25.34%
34.81%
69.28%
75.95%
40.27%
34.54%
60.13%
3.87%
37.55%
33.59%
32.17%
34.00%
65.18%
60.23%
58.38%
28.50%
29.46%
Llama-3.1-8B
65.10%
43.21%
55.73%
78.95%
81.01%
53.41%
61.59%
77.72%
26.00%
43.01%
52.29%
63.84%
60.00%
57.51%
55.28%
53.81%
41.44%
38.39%
Qwen2.5-7B
74.18%
51.77%
65.08%
78.95%
79.71%
51.37%
71.72%
80.37%
9.40%
48.66%
59.40%
76.81%
65.70%
59.68%
57.51%
55.44%
47.33%
49.26%
Citation
If you use Fanar-1-9B or
Fanar-1-9B-Instruct
or the Fanar GenAI system in your research or applications, please cite:
@misc{fanarllm2025,
title={Fanar: An Arabic-Centric Multimodal Generative AI Platform},
author={Fanar Team and Ummar Abbas and Mohammad Shahmeer Ahmad and Firoj Alam and Enes Altinisik and Ehsannedin Asgari and Yazan Boshmaf and Sabri Boughorbel and Sanjay Chawla and Shammur Chowdhury and Fahim Dalvi and Kareem Darwish and Nadir Durrani and Mohamed Elfeky and Ahmed Elmagarmid and Mohamed Eltabakh and Masoomali Fatehkia and Anastasios Fragkopoulos and Maram Hasanain and Majd Hawasly and Mus'ab Husaini and Soon-Gyo Jung and Ji Kim Lucas and Walid Magdy and Safa Messaoud and Abubakr Mohamed and Tasnim Mohiuddin and Basel Mousi and Hamdy Mubarak and Ahmad Musleh and Zan Naeem and Mourad Ouzzani and Dorde Popovic and Amin Sadeghi and Husrev Taha Sencar and Mohammed Shinoy and Omar Sinan and Yifan Zhang and Ahmed Ali and Yassine El Kheir and Xiaosong Ma and Chaoyi Ruan}},
year={2025},
url={https://arxiv.org/abs/2501.13944},
}
Fanar-1-9B huggingface.co is an AI model on huggingface.co that provides Fanar-1-9B's model effect (), which can be used instantly with this QCRI Fanar-1-9B model. huggingface.co supports a free trial of the Fanar-1-9B model, and also provides paid use of the Fanar-1-9B. Support call Fanar-1-9B model through api, including Node.js, Python, http.
Fanar-1-9B huggingface.co is an online trial and call api platform, which integrates Fanar-1-9B's modeling effects, including api services, and provides a free online trial of Fanar-1-9B, you can try Fanar-1-9B online for free by clicking the link below.
QCRI Fanar-1-9B online free url in huggingface.co:
Fanar-1-9B is an open source model from GitHub that offers a free installation service, and any user can find Fanar-1-9B on GitHub to install. At the same time, huggingface.co provides the effect of Fanar-1-9B install, users can directly use Fanar-1-9B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.