Tigrinya Abusive Language Detection (TiALD) Dataset
is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of
13,717 YouTube comments
annotated for
abusiveness
,
sentiment
, and
topic
tasks. The dataset includes comments written in both the
Ge’ez script
and prevalent non-standard Latin
transliterations
to mirror real-world usage.
⚠️ The dataset contains explicit, obscene, and potentially hateful language. It should be used for research purposes only. ⚠️
The following hyperparameters were used during training:
learning_rate: 2e-05
train_batch_size: 16
optimizer: Adam (betas=0.9, 0.999, epsilon=1e-08)
lr_scheduler_type: linear
num_epochs: 4.0
seed: 42
Intended Usage
The TiALD dataset and models designed to support:
Research in abusive language detection in low-resource languages
Context-aware abuse, sentiment, and topic modeling
Multi-task and transfer learning with digraphic scripts
Evaluation of multilingual and fine-tuned language models
Researchers and developers should avoid using this dataset for direct moderation or enforcement tasks without human oversight.
Ethical Considerations
Sensitive content
: Contains toxic and offensive language. Use for research purposes only.
Cultural sensitivity
: Abuse is context-dependent; annotations were made by native speakers to account for cultural nuance.
Bias mitigation
: Data sampling and annotation were carefully designed to minimize reinforcement of stereotypes.
Privacy
: All the source content for the dataset is publicly available on YouTube.
Respect for expression
: The dataset should not be used for automated censorship without human review.
This research received IRB approval (Ref: KH2022-133) and followed ethical data collection and annotation practices, including informed consent of annotators.
Citation
If you use this model or the
TiALD
dataset in your work, please cite:
@misc{gaim-etal-2025-tiald-benchmark,
title = {A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings},
author = {Fitsum Gaim and Hoyun Song and Huije Lee and Changgeon Ko and Eui Jun Hwang and Jong C. Park},
year = {2025},
eprint = {2505.12116},
archiveprefix = {arXiv},
primaryclass = {cs.CL},
url = {https://arxiv.org/abs/2505.12116}
}
tiroberta-abusiveness-detection huggingface.co is an AI model on huggingface.co that provides tiroberta-abusiveness-detection's model effect (), which can be used instantly with this fgaim tiroberta-abusiveness-detection model. huggingface.co supports a free trial of the tiroberta-abusiveness-detection model, and also provides paid use of the tiroberta-abusiveness-detection. Support call tiroberta-abusiveness-detection model through api, including Node.js, Python, http.
tiroberta-abusiveness-detection huggingface.co is an online trial and call api platform, which integrates tiroberta-abusiveness-detection's modeling effects, including api services, and provides a free online trial of tiroberta-abusiveness-detection, you can try tiroberta-abusiveness-detection online for free by clicking the link below.
fgaim tiroberta-abusiveness-detection online free url in huggingface.co:
tiroberta-abusiveness-detection is an open source model from GitHub that offers a free installation service, and any user can find tiroberta-abusiveness-detection on GitHub to install. At the same time, huggingface.co provides the effect of tiroberta-abusiveness-detection install, users can directly use tiroberta-abusiveness-detection installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
tiroberta-abusiveness-detection install url in huggingface.co: