This model is a fine-tuned version of the
DistilBERT model
to classify toxic comments.
How to use
You can use the model with the following code.
from transformers import AutoModelForSequenceClassification, AutoTokenizer, TextClassificationPipeline
model_path = "martin-ha/toxic-comment-model"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForSequenceClassification.from_pretrained(model_path)
pipeline = TextClassificationPipeline(model=model, tokenizer=tokenizer)
print(pipeline('This is a test text.'))
Limitations and Bias
This model is intended to use for classify toxic online classifications. However, one limitation of the model is that it performs poorly for some comments that mention a specific identity subgroup, like Muslim. The following table shows a evaluation score for different identity group. You can learn the specific meaning of this metrics
here
. But basically, those metrics shows how well a model performs for a specific group. The larger the number, the better.
subgroup
subgroup_size
subgroup_auc
bpsn_auc
bnsp_auc
muslim
108
0.689
0.811
0.88
jewish
40
0.749
0.86
0.825
homosexual_gay_or_lesbian
56
0.795
0.706
0.972
black
84
0.866
0.758
0.975
white
112
0.876
0.784
0.97
female
306
0.898
0.887
0.948
christian
231
0.904
0.917
0.93
male
225
0.922
0.862
0.967
psychiatric_or_mental_illness
26
0.924
0.907
0.95
The table above shows that the model performs poorly for the muslim and jewish group. In fact, you pass the sentence "Muslims are people who follow or practice Islam, an Abrahamic monotheistic religion." Into the model, the model will classify it as toxic. Be mindful for this type of potential bias.
Training data
The training data comes this
Kaggle competition
. We use 10% of the
train.csv
data to train the model.
The model achieves 94% accuracy and 0.59 f1-score in a 10000 rows held-out test set.
Runs of garak-llm toxic-comment-model on huggingface.co
30.6K
Total runs
0
24-hour runs
-1.9K
3-day runs
788
7-day runs
-10.8K
30-day runs
More Information About toxic-comment-model huggingface.co Model
toxic-comment-model huggingface.co
toxic-comment-model huggingface.co is an AI model on huggingface.co that provides toxic-comment-model's model effect (), which can be used instantly with this garak-llm toxic-comment-model model. huggingface.co supports a free trial of the toxic-comment-model model, and also provides paid use of the toxic-comment-model. Support call toxic-comment-model model through api, including Node.js, Python, http.
toxic-comment-model huggingface.co is an online trial and call api platform, which integrates toxic-comment-model's modeling effects, including api services, and provides a free online trial of toxic-comment-model, you can try toxic-comment-model online for free by clicking the link below.
garak-llm toxic-comment-model online free url in huggingface.co:
toxic-comment-model is an open source model from GitHub that offers a free installation service, and any user can find toxic-comment-model on GitHub to install. At the same time, huggingface.co provides the effect of toxic-comment-model install, users can directly use toxic-comment-model installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
toxic-comment-model install url in huggingface.co: