Text classification model based on
classla/bcms-bertic
and fine-tuned on the
FRENK dataset
comprising of LGBT and migrant hatespeech. Only the Croatian subset of the data was used for fine-tuning and the dataset has been relabeled for binary classification (offensive or acceptable).
Fine-tuning hyperparameters
Fine-tuning was performed with
simpletransformers
. Beforehand a brief hyperparameter optimisation was performed and the presumed optimal hyperparameters are:
The same pipeline was run with two other transformer models and
fasttext
for comparison. Accuracy and macro F1 score were recorded for each of the 6 fine-tuning sessions and post festum analyzed.
model
average accuracy
average macro F1
bcms-bertic-frenk-hate
0.8313
0.8219
EMBEDDIA/crosloengual-bert
0.8054
0.796
xlm-roberta-base
0.7175
0.7049
fasttext
0.771
0.754
From recorded accuracies and macro F1 scores p-values were also calculated:
Comparison with
crosloengual-bert
:
test
accuracy p-value
macro F1 p-value
Wilcoxon
0.00781
0.00781
Mann Whithney
0.00108
0.00108
Student t-test
2.43e-10
1.27e-10
Comparison with
xlm-roberta-base
:
test
accuracy p-value
macro F1 p-value
Wilcoxon
0.00781
0.00781
Mann Whithney
0.00107
0.00108
Student t-test
4.83e-11
5.61e-11
Use examples
from simpletransformers.classification import ClassificationModel
model = ClassificationModel(
"bert", "5roop/bcms-bertic-frenk-hate", use_cuda=True,
)
predictions, logit_output = model.predict(['Ne odbacujem da će RH primiti još migranata iz Afganistana, no neće biti novog vala',
"Potpredsjednik Vlade i ministar branitelja Tomo Medved komentirao je Vladine planove za zakonsku zabranu pozdrava 'za dom spremni' "])
predictions
### Output:### array([0, 0])
Citation
If you use the model, please cite the following paper on which the original model is based:
@inproceedings{ljubesic-lauc-2021-bertic,
title = "{BERT}i{\'c} - The Transformer Language Model for {B}osnian, {C}roatian, {M}ontenegrin and {S}erbian",
author = "Ljube{\v{s}}i{\'c}, Nikola and Lauc, Davor",
booktitle = "Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing",
month = apr,
year = "2021",
address = "Kiyv, Ukraine",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/2021.bsnlp-1.5",
pages = "37--42",
}
and the dataset used for fine-tuning:
@misc{ljubešić2019frenk,
title={The FRENK Datasets of Socially Unacceptable Discourse in Slovene and English},
author={Nikola Ljubešić and Darja Fišer and Tomaž Erjavec},
year={2019},
eprint={1906.02045},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1906.02045}
}
Runs of classla bcms-bertic-frenk-hate on huggingface.co
38
Total runs
0
24-hour runs
8
3-day runs
14
7-day runs
19
30-day runs
More Information About bcms-bertic-frenk-hate huggingface.co Model
bcms-bertic-frenk-hate huggingface.co is an AI model on huggingface.co that provides bcms-bertic-frenk-hate's model effect (), which can be used instantly with this classla bcms-bertic-frenk-hate model. huggingface.co supports a free trial of the bcms-bertic-frenk-hate model, and also provides paid use of the bcms-bertic-frenk-hate. Support call bcms-bertic-frenk-hate model through api, including Node.js, Python, http.
bcms-bertic-frenk-hate huggingface.co is an online trial and call api platform, which integrates bcms-bertic-frenk-hate's modeling effects, including api services, and provides a free online trial of bcms-bertic-frenk-hate, you can try bcms-bertic-frenk-hate online for free by clicking the link below.
classla bcms-bertic-frenk-hate online free url in huggingface.co:
bcms-bertic-frenk-hate is an open source model from GitHub that offers a free installation service, and any user can find bcms-bertic-frenk-hate on GitHub to install. At the same time, huggingface.co provides the effect of bcms-bertic-frenk-hate install, users can directly use bcms-bertic-frenk-hate installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
bcms-bertic-frenk-hate install url in huggingface.co: