A lightweight Arabic dialect identification model that classifies input text into one of
11 Arabic dialect / language codes
. It is used as the routing backbone in the
Lahgtna
pipeline to automatically select the correct voice reference and Chatterbox language token for speech synthesis.
Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.
Secondary use
Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.
Out-of-scope use
Non-Arabic languages
Code-switched text (Arabic + English mixed)
Dialect intensity scoring or fine-grained subdialect classification
High-stakes decisions without human review
How to Use
Direct inference
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "oddadmix/dialect-router-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect) # e.g. "eg"
With the Transformers pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="oddadmix/dialect-router-v0.1",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]
Inside Lahgtna TTS
from inference import run_pipeline
# Dialect is detected automatically
run_pipeline(
text="اه ياراسي الواحد دماغه وجعاه",
output_path="output.wav",
)
Limitations & Biases
Short texts
(< 5 tokens) may produce unreliable predictions — the model benefits from sentence-length input.
Code-switched text
(e.g. Arabic + French in Moroccan Darija, or Arabic + English) may confuse the classifier.
Dialect continuum
— dialects from geographically adjacent regions (e.g.
sy
/
lb
,
eg
/
ly
) may be confused by the model.
Corpus bias
— label distribution in training data may not reflect real-world dialect prevalence; some dialects (e.g.
sd
,
ly
) may have lower recall.
This model should
not
be used for identity classification of individuals.
Citation
If you use this model in your research or product, please cite:
@misc{lahgtna-dialect-router-2025,
title = {dialect-router-v0.1: Arabic Dialect Identification for TTS Routing},
author = {Oddadmix},
year = {2025},
url = {https://huggingface.co/oddadmix/dialect-router-v0.1}
}
dialect-router-v0.1 huggingface.co is an AI model on huggingface.co that provides dialect-router-v0.1's model effect (), which can be used instantly with this oddadmix dialect-router-v0.1 model. huggingface.co supports a free trial of the dialect-router-v0.1 model, and also provides paid use of the dialect-router-v0.1. Support call dialect-router-v0.1 model through api, including Node.js, Python, http.
dialect-router-v0.1 huggingface.co is an online trial and call api platform, which integrates dialect-router-v0.1's modeling effects, including api services, and provides a free online trial of dialect-router-v0.1, you can try dialect-router-v0.1 online for free by clicking the link below.
oddadmix dialect-router-v0.1 online free url in huggingface.co:
dialect-router-v0.1 is an open source model from GitHub that offers a free installation service, and any user can find dialect-router-v0.1 on GitHub to install. At the same time, huggingface.co provides the effect of dialect-router-v0.1 install, users can directly use dialect-router-v0.1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
dialect-router-v0.1 install url in huggingface.co: