SPhoBERT fine-tuned cho bài toán phân loại Hate Speech tiếng Việt. Model này được fine-tune từ
vinai/phobert-base
trên dataset ViHSD (Vietnamese Hate Speech Dataset).
How to Use
Basic Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Load model and tokenizer
model_name = "visolex/hate-speech-sphobert"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Classify text
text = "Văn bản tiếng Việt cần phân loại"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_label = torch.argmax(predictions, dim=-1).item()
# Label mapping
label_names = {
0: "CLEAN",
1: "OFFENSIVE",
2: "HATE"
}
print(f"Predicted label: {label_names[predicted_label]}")
print(f"Confidence scores: {predictions[0].tolist()}")
Using the Pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="visolex/hate-speech-sphobert",
tokenizer="visolex/hate-speech-sphobert"
)
result = classifier("Văn bản tiếng Việt cần phân loại")
print(result)
Training Details
Training Data
Dataset: ViHSD (Vietnamese Hate Speech Dataset)
Training samples: ~8,000 samples
Validation samples: ~1,000 samples
Test samples: ~1,000 samples
Training Procedure
Framework: PyTorch + Transformers
Optimizer: AdamW
Learning Rate: 2e-5
Batch Size: 32
Epochs: Varies by model
Max Sequence Length: 256
Label Distribution
CLEAN (0): Normal content without offensive language
OFFENSIVE (1): Mildly offensive content
HATE (2): Hate speech and extremist language
Evaluation
Model được đánh giá trên test set của ViHSD với các metrics:
Accuracy: Overall classification accuracy
F1 Macro: Macro-averaged F1 score across all labels
F1 Weighted: Weighted F1 score based on label frequency
Limitations and Bias
Model chỉ được train trên dữ liệu tiếng Việt từ mạng xã hội
Performance có thể giảm trên domain khác (email, document, etc.)
Model có thể có bias từ dữ liệu training
Cần đánh giá thêm trên dữ liệu real-world
Citation
Contact
License
This model is distributed under the MIT License.
Runs of visolex sphobert-hsd on huggingface.co
18
Total runs
0
24-hour runs
1
3-day runs
-1
7-day runs
-18
30-day runs
More Information About sphobert-hsd huggingface.co Model
sphobert-hsd huggingface.co is an AI model on huggingface.co that provides sphobert-hsd's model effect (), which can be used instantly with this visolex sphobert-hsd model. huggingface.co supports a free trial of the sphobert-hsd model, and also provides paid use of the sphobert-hsd. Support call sphobert-hsd model through api, including Node.js, Python, http.
sphobert-hsd huggingface.co is an online trial and call api platform, which integrates sphobert-hsd's modeling effects, including api services, and provides a free online trial of sphobert-hsd, you can try sphobert-hsd online for free by clicking the link below.
visolex sphobert-hsd online free url in huggingface.co:
sphobert-hsd is an open source model from GitHub that offers a free installation service, and any user can find sphobert-hsd on GitHub to install. At the same time, huggingface.co provides the effect of sphobert-hsd install, users can directly use sphobert-hsd installed effect in huggingface.co for debugging and trial. It also supports api for free installation.