A lightweight Arabic dialect identification model that classifies input text into one of
15 language codes: 13 Arabic dialects, Modern Standard Arabic, and English
. It is the routing backbone in the
Lahgtna
pipeline, automatically selecting the correct voice reference and Chatterbox language token for speech synthesis.
v0.2 is a fine-tune of
asafaya/bert-mini-arabic
and expands coverage from 10 to 13 Arabic dialects, and adds an English label.
Model Details
Property
Value
Base model
asafaya/bert-mini-arabic
Architecture
BERT-mini encoder + sequence classification head
Task
Multi-class text classification (15 classes)
Input
Raw text (up to 512 tokens)
Output
One of 15 dialect / language codes
Languages
Arabic (ar), English (en)
License
MIT
Evaluation Results
Metric
Score
Accuracy
0.9359
F1 Macro
0.9052
Eval Loss
0.4537
Dialect Labels
ID
Label
Dialect / Language
Region
0
ar
Modern Standard Arabic (MSA)
—
1
bh
Bahraini
Bahrain
2
dz
Algerian
Algeria
3
eg
Egyptian
Egypt
4
en
English
—
5
iq
Iraqi
Iraq
6
lb
Lebanese
Lebanon
7
ly
Libyan
Libya
8
ma
Moroccan (Darija)
Morocco
9
ps
Palestinian
Palestine
10
sa
Saudi
Saudi Arabia
11
sd
Sudanese
Sudan
12
sy
Syrian
Syria
13
tn
Tunisian
Tunisia
14
ye
Yemeni
Yemen
What's New in v0.2
13 Arabic dialects
(up from 10): adds Bahraini (
bh
), Algerian (
dz
), and Yemeni (
ye
)
English label
(
en
) — English input is now routed explicitly instead of being out-of-scope
Moroccan label renamed
mo
→
ma
(ISO 3166 country code)
New base model:
asafaya/bert-mini-arabic
— smaller and faster for routing workloads
Retrained on an expanded multi-dialect corpus
Intended Use
Primary use
Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.
Secondary use
Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.
Out-of-scope use
Languages other than Arabic and English
Code-switched text (Arabic + English mixed)
Dialect intensity scoring or fine-grained subdialect classification
High-stakes decisions without human review
How to Use
Direct inference
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "oddadmix/dialect-router-v0.2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect) # e.g. "eg"
With the Transformers pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="oddadmix/dialect-router-v0.2",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]
Inside Lahgtna TTS
from inference import run_pipeline
# Dialect is detected automatically
run_pipeline(
text="اه ياراسي الواحد دماغه وجعاه",
output_path="output.wav",
)
Short texts
(< 5 tokens) may produce unreliable predictions — the model benefits from sentence-length input.
Code-switched text
(e.g. Arabic + French in Maghrebi dialects, or Arabic + English) may confuse the classifier; heavily mixed input may be routed to
en
.
Dialect continuum
— dialects from geographically adjacent regions (e.g. sy / lb / ps, ma / dz / tn, sa / bh) may be confused by the model.
Corpus bias
— label distribution in training data may not reflect real-world dialect prevalence; some dialects (e.g. sd, ly, bh, ye) may have lower recall.
This model should not be used for identity classification of individuals.
Citation
@misc{lahgtna-dialect-router-2026,
title = {dialect-router-v0.2: Arabic Dialect Identification for TTS Routing},
author = {Oddadmix},
year = {2026},
url = {https://huggingface.co/oddadmix/dialect-router-v0.2}
}
Runs of oddadmix dialect-router-v0.2 on huggingface.co
161
Total runs
0
24-hour runs
53
3-day runs
89
7-day runs
153
30-day runs
More Information About dialect-router-v0.2 huggingface.co Model
dialect-router-v0.2 huggingface.co
dialect-router-v0.2 huggingface.co is an AI model on huggingface.co that provides dialect-router-v0.2's model effect (), which can be used instantly with this oddadmix dialect-router-v0.2 model. huggingface.co supports a free trial of the dialect-router-v0.2 model, and also provides paid use of the dialect-router-v0.2. Support call dialect-router-v0.2 model through api, including Node.js, Python, http.
dialect-router-v0.2 huggingface.co is an online trial and call api platform, which integrates dialect-router-v0.2's modeling effects, including api services, and provides a free online trial of dialect-router-v0.2, you can try dialect-router-v0.2 online for free by clicking the link below.
oddadmix dialect-router-v0.2 online free url in huggingface.co:
dialect-router-v0.2 is an open source model from GitHub that offers a free installation service, and any user can find dialect-router-v0.2 on GitHub to install. At the same time, huggingface.co provides the effect of dialect-router-v0.2 install, users can directly use dialect-router-v0.2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
dialect-router-v0.2 install url in huggingface.co: