This is a roBERTa-base model trained on ~58M tweets and finetuned for hate speech detection with the TweetEval benchmark.
This model is specialized to detect hate speech against women and immigrants.
from transformers import AutoModelForSequenceClassification
from transformers import TFAutoModelForSequenceClassification
from transformers import AutoTokenizer
import numpy as np
from scipy.special import softmax
import csv
import urllib.request
# Preprocess text (username and link placeholders)defpreprocess(text):
new_text = []
for t in text.split(" "):
t = '@user'if t.startswith('@') andlen(t) > 1else t
t = 'http'if t.startswith('http') else t
new_text.append(t)
return" ".join(new_text)
# Tasks:# emoji, emotion, hate, irony, offensive, sentiment# stance/abortion, stance/atheism, stance/climate, stance/feminist, stance/hillary
task='hate'
MODEL = f"cardiffnlp/twitter-roberta-base-{task}"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
# download label mapping
labels=[]
mapping_link = f"https://raw.githubusercontent.com/cardiffnlp/tweeteval/main/datasets/{task}/mapping.txt"with urllib.request.urlopen(mapping_link) as f:
html = f.read().decode('utf-8').split("\n")
csvreader = csv.reader(html, delimiter='\t')
labels = [row[1] for row in csvreader iflen(row) > 1]
# PT
model = AutoModelForSequenceClassification.from_pretrained(MODEL)
model.save_pretrained(MODEL)
text = "Good night 😊"
text = preprocess(text)
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
scores = output[0][0].detach().numpy()
scores = softmax(scores)
# # TF# model = TFAutoModelForSequenceClassification.from_pretrained(MODEL)# model.save_pretrained(MODEL)# text = "Good night 😊"# encoded_input = tokenizer(text, return_tensors='tf')# output = model(encoded_input)# scores = output[0][0].numpy()# scores = softmax(scores)
ranking = np.argsort(scores)
ranking = ranking[::-1]
for i inrange(scores.shape[0]):
l = labels[ranking[i]]
s = scores[ranking[i]]
print(f"{i+1}) {l}{np.round(float(s), 4)}")
Output:
1) not-hate 0.9168
2) hate 0.0832
Runs of cardiffnlp twitter-roberta-base-hate on huggingface.co
1.2K
Total runs
0
24-hour runs
-53
3-day runs
407
7-day runs
-239
30-day runs
More Information About twitter-roberta-base-hate huggingface.co Model
twitter-roberta-base-hate huggingface.co
twitter-roberta-base-hate huggingface.co is an AI model on huggingface.co that provides twitter-roberta-base-hate's model effect (), which can be used instantly with this cardiffnlp twitter-roberta-base-hate model. huggingface.co supports a free trial of the twitter-roberta-base-hate model, and also provides paid use of the twitter-roberta-base-hate. Support call twitter-roberta-base-hate model through api, including Node.js, Python, http.
twitter-roberta-base-hate huggingface.co is an online trial and call api platform, which integrates twitter-roberta-base-hate's modeling effects, including api services, and provides a free online trial of twitter-roberta-base-hate, you can try twitter-roberta-base-hate online for free by clicking the link below.
cardiffnlp twitter-roberta-base-hate online free url in huggingface.co:
twitter-roberta-base-hate is an open source model from GitHub that offers a free installation service, and any user can find twitter-roberta-base-hate on GitHub to install. At the same time, huggingface.co provides the effect of twitter-roberta-base-hate install, users can directly use twitter-roberta-base-hate installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
twitter-roberta-base-hate install url in huggingface.co: