PhoBERT Base fine-tuned cho bài toán phân loại Hate Speech tiếng Việt. Model này được fine-tune từ
vinai/phobert-base
trên dataset ViHSD (Vietnamese Hate Speech Dataset).
How to Use
Basic Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Load model and tokenizer
model_name = "visolex/hate-speech-phobert-v1"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
# Classify text
text = "Văn bản tiếng Việt cần phân loại"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_label = torch.argmax(predictions, dim=-1).item()
# Label mapping
label_names = {
0: "CLEAN",
1: "OFFENSIVE",
2: "HATE"
}
print(f"Predicted label: {label_names[predicted_label]}")
print(f"Confidence scores: {predictions[0].tolist()}")
Using the Pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="visolex/hate-speech-phobert-v1",
tokenizer="visolex/hate-speech-phobert-v1"
)
result = classifier("Văn bản tiếng Việt cần phân loại")
print(result)
Training Details
Training Data
Dataset: ViHSD (Vietnamese Hate Speech Dataset)
Training samples: ~8,000 samples
Validation samples: ~1,000 samples
Test samples: ~1,000 samples
Training Procedure
Framework: PyTorch + Transformers
Optimizer: AdamW
Learning Rate: 2e-5
Batch Size: 32
Epochs: Varies by model
Max Sequence Length: 256
Label Distribution
CLEAN (0): Normal content without offensive language
OFFENSIVE (1): Mildly offensive content
HATE (2): Hate speech and extremist language
Evaluation
Model được đánh giá trên test set của ViHSD với các metrics:
Accuracy: Overall classification accuracy
F1 Macro: Macro-averaged F1 score across all labels
F1 Weighted: Weighted F1 score based on label frequency
Limitations and Bias
Model chỉ được train trên dữ liệu tiếng Việt từ mạng xã hội
Performance có thể giảm trên domain khác (email, document, etc.)
Model có thể có bias từ dữ liệu training
Cần đánh giá thêm trên dữ liệu real-world
Citation
Contact
License
This model is distributed under the MIT License.
Runs of visolex phobert-v1-hsd on huggingface.co
35
Total runs
0
24-hour runs
7
3-day runs
9
7-day runs
15
30-day runs
More Information About phobert-v1-hsd huggingface.co Model
phobert-v1-hsd huggingface.co is an AI model on huggingface.co that provides phobert-v1-hsd's model effect (), which can be used instantly with this visolex phobert-v1-hsd model. huggingface.co supports a free trial of the phobert-v1-hsd model, and also provides paid use of the phobert-v1-hsd. Support call phobert-v1-hsd model through api, including Node.js, Python, http.
phobert-v1-hsd huggingface.co is an online trial and call api platform, which integrates phobert-v1-hsd's modeling effects, including api services, and provides a free online trial of phobert-v1-hsd, you can try phobert-v1-hsd online for free by clicking the link below.
visolex phobert-v1-hsd online free url in huggingface.co:
phobert-v1-hsd is an open source model from GitHub that offers a free installation service, and any user can find phobert-v1-hsd on GitHub to install. At the same time, huggingface.co provides the effect of phobert-v1-hsd install, users can directly use phobert-v1-hsd installed effect in huggingface.co for debugging and trial. It also supports api for free installation.