This model is trained on the dataset of sensitive topics of the Russian language. The concept of sensitive topics is described
in this article
presented at the workshop for Balto-Slavic NLP at the EACL-2021 conference. Please note that this article describes the first version of the dataset, while the model is trained on the extended version of the dataset open-sourced on our
GitHub
or on
kaggle
. The properties of the dataset is the same as the one described in the article, the only difference is the size.
Instructions
The model predicts combinations of 18 sensitive topics described in the
article
. You can find step-by-step instructions for using the model
here
Metrics
The dataset partially manually labeled samples and partially semi-automatically labeled samples. Learn more in our article. We tested the performance of the classifier only on the part of manually labeled data that is why some topics are not well represented in the test set.
If you find this repository helpful, feel free to cite our publication:
@inproceedings{babakov-etal-2021-detecting,
title = "Detecting Inappropriate Messages on Sensitive Topics that Could Harm a Company{'}s Reputation",
author = "Babakov, Nikolay and
Logacheva, Varvara and
Kozlova, Olga and
Semenov, Nikita and
Panchenko, Alexander",
booktitle = "Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing",
month = apr,
year = "2021",
address = "Kiyv, Ukraine",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/2021.bsnlp-1.4",
pages = "26--36",
abstract = "Not all topics are equally {``}flammable{''} in terms of toxicity: a calm discussion of turtles or fishing less often fuels inappropriate toxic dialogues than a discussion of politics or sexual minorities. We define a set of sensitive topics that can yield inappropriate and toxic messages and describe the methodology of collecting and labelling a dataset for appropriateness. While toxicity in user-generated data is well-studied, we aim at defining a more fine-grained notion of inappropriateness. The core of inappropriateness is that it can harm the reputation of a speaker. This is different from toxicity in two respects: (i) inappropriateness is topic-related, and (ii) inappropriate message is not toxic but still unacceptable. We collect and release two datasets for Russian: a topic-labelled dataset and an appropriateness-labelled dataset. We also release pre-trained classification models trained on this data.",
}
Runs of apanc russian-sensitive-topics on huggingface.co
281
Total runs
27
24-hour runs
74
3-day runs
28
7-day runs
-254
30-day runs
More Information About russian-sensitive-topics huggingface.co Model
russian-sensitive-topics huggingface.co
russian-sensitive-topics huggingface.co is an AI model on huggingface.co that provides russian-sensitive-topics's model effect (), which can be used instantly with this apanc russian-sensitive-topics model. huggingface.co supports a free trial of the russian-sensitive-topics model, and also provides paid use of the russian-sensitive-topics. Support call russian-sensitive-topics model through api, including Node.js, Python, http.
russian-sensitive-topics huggingface.co is an online trial and call api platform, which integrates russian-sensitive-topics's modeling effects, including api services, and provides a free online trial of russian-sensitive-topics, you can try russian-sensitive-topics online for free by clicking the link below.
apanc russian-sensitive-topics online free url in huggingface.co:
russian-sensitive-topics is an open source model from GitHub that offers a free installation service, and any user can find russian-sensitive-topics on GitHub to install. At the same time, huggingface.co provides the effect of russian-sensitive-topics install, users can directly use russian-sensitive-topics installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
russian-sensitive-topics install url in huggingface.co: