SA-BERT-V1 delivers unparalleled Saudi-dialect understanding—achieving a +0.0022 in-vs-cross similarity gap and 0.98 mean cosine scores across 44 specialized categories, setting a new standard for Arabic dialect sentence embeddings.”
▪️
SA-BERT-V1
shows a positive in–cross gap and high absolute similarity, proving the effectiveness of targeted Saudi-dialect fine-tuning.
▪️
In vs Cross:
Both near ~0.98, with a slight positive gap (+0.0023), meaning same-topic embeddings are closer.
▪️
Performance:
Exceptional clustering for Saudi dialect; ideal for retrieval or grouping tasks.
▪️The evaluations—both the similarity metrics and the “in- vs-cross” gap plots—were run on a held-out test set of
1280 Saudi-dialect sentences covering 44 diverse categories
(e.g. Greetings, Weather, Law & Justice, etc.).
▪️
Dataset
is create by the space and released to evaluate embedding models by sampling intra-category and cross-category pairs from that set to compute:
◽️Average in-category / cross-category cosine similarities ◽️Top-5 most/least similar pairs ◽️Per-category average similarities
import torch
from transformers import AutoTokenizer, AutoModel
# Configuration
MODEL_ID = "Omartificial-Intelligence-Space/SA-BERT-V1"
DEVICE = torch.device("cuda"if torch.cuda.is_available() else"cpu")
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID , token= "PASS_READ_TOKEN_HERE")
model = AutoModel.from_pretrained(MODEL_ID , token = "PASS_READ_TOKEN_HERE").to(DEVICE).eval()
defembed_sentence(text: str) -> torch.Tensor:
""" Tokenizes `text`, feeds it through SA-BERT-V1, and returns a 768-dimensional mean-pooled sentence embedding. """# Encode the text
enc = tokenizer(
text,
truncation=True,
padding="max_length",
max_length=256,
return_tensors="pt"
).to(DEVICE)
# Forward passwith torch.no_grad():
outputs = model(**enc).last_hidden_state # shape: (1, seq_len, 768)# Mean-pooling over valid tokens
mask = enc["attention_mask"].unsqueeze(-1) # shape: (1, seq_len, 1)
summed = (outputs * mask).sum(dim=1) # shape: (1, 768)
counts = mask.sum(dim=1).clamp(min=1e-9) # shape: (1, 1)
embedding = summed / counts # shape: (1, 768)return embedding.squeeze(0) # shape: (768,)# Example usageif __name__ == "__main__":
sentences = [
"شتبي من البقالة؟",
"كيف حالك؟",
"وش رايك في الموضوع هذا؟"
]
for s in sentences:
vec = embed_sentence(s)
print(f"Sentence: {s}\nEmbedding shape: {vec.shape}\n")
Citation
If you use MarBERTv2-SA in your research or applications, please cite:
@misc{nacar2025SABERTV1,
title={SA-BERT-V1: Fine-Tuned Saudi-Dialect Embeddings},
author={NAcar, Omer},
year={2025},
publisher={Omartificial-Intelligence-Space},
howpublished={\url{https://huggingface.co/Omartificial-Intelligence-Space/SA-BERT-V1}},
}
@inproceedings{abdul-mageed-etal-2021-arbert,
title = "{ARBERT} {\&} {MARBERT}: Deep Bidirectional Transformers for {A}rabic",
author = "Abdul-Mageed, Muhammad and Elmadany, AbdelRahim and Nagoudi, El Moatez Billah",
booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
year = "2021",
publisher = "Association for Computational Linguistics",
pages = "7088--7105",
}
Runs of Omartificial-Intelligence-Space SA-BERT-V1 on huggingface.co
18
Total runs
0
24-hour runs
-2
3-day runs
-2
7-day runs
11
30-day runs
More Information About SA-BERT-V1 huggingface.co Model
SA-BERT-V1 huggingface.co is an AI model on huggingface.co that provides SA-BERT-V1's model effect (), which can be used instantly with this Omartificial-Intelligence-Space SA-BERT-V1 model. huggingface.co supports a free trial of the SA-BERT-V1 model, and also provides paid use of the SA-BERT-V1. Support call SA-BERT-V1 model through api, including Node.js, Python, http.
SA-BERT-V1 huggingface.co is an online trial and call api platform, which integrates SA-BERT-V1's modeling effects, including api services, and provides a free online trial of SA-BERT-V1, you can try SA-BERT-V1 online for free by clicking the link below.
Omartificial-Intelligence-Space SA-BERT-V1 online free url in huggingface.co:
SA-BERT-V1 is an open source model from GitHub that offers a free installation service, and any user can find SA-BERT-V1 on GitHub to install. At the same time, huggingface.co provides the effect of SA-BERT-V1 install, users can directly use SA-BERT-V1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.