A binary safety classifier for detecting unsafe content in LLM inputs. Part of a hierarchical MLCommons-aligned safety classification system.
Model Description
This is
Level 1
of a 2-level hierarchical safety classifier:
Level 1 (this model)
: Binary classification (safe/unsafe) - high recall for catching threats
Level 2
: 9-class hazard taxonomy (MLCommons AI Safety aligned) - for categorizing unsafe content
Performance
Metric
Score
Accuracy
84.9%
F1 Score
84.9%
Labels
ID
Label
Description
0
safe
Content is safe
1
unsafe
Content is potentially harmful
Training Hyperparameters
Parameter
Value
Base Model
jhu-clsp/mmBERT-base
LoRA Rank
32
LoRA Alpha
64
LoRA Dropout
0.1
Learning Rate
5e-5
Epochs
10
Batch Size
64
Max Samples
18,000
Training Samples
12,600
Validation Samples
2,700
Usage
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel
# Load model
tokenizer = AutoTokenizer.from_pretrained("llm-semantic-router/mmbert-safety-classifier-level1")
base_model = AutoModelForSequenceClassification.from_pretrained(
"jhu-clsp/mmBERT-base",
num_labels=2,
torch_dtype=torch.float32
)
model = PeftModel.from_pretrained(base_model, "llm-semantic-router/mmbert-safety-classifier-level1")
model.eval()
# Inference
text = "How do I bake a chocolate cake?"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
pred = outputs.logits.argmax(-1).item()
labels = {0: "safe", 1: "unsafe"}
print(f"Prediction: {labels[pred]}")
Hierarchical Classification Pipeline
For complete safety classification, use with Level 2:
# If Level 1 predicts "unsafe", run Level 2 for hazard categoryif pred == 1: # unsafe# Load and run Level 2 model for specific hazard category# See: llm-semantic-router/mmbert-safety-classifier-level2pass
mmbert-safety-classifier-level1 huggingface.co is an AI model on huggingface.co that provides mmbert-safety-classifier-level1's model effect (), which can be used instantly with this llm-semantic-router mmbert-safety-classifier-level1 model. huggingface.co supports a free trial of the mmbert-safety-classifier-level1 model, and also provides paid use of the mmbert-safety-classifier-level1. Support call mmbert-safety-classifier-level1 model through api, including Node.js, Python, http.
mmbert-safety-classifier-level1 huggingface.co is an online trial and call api platform, which integrates mmbert-safety-classifier-level1's modeling effects, including api services, and provides a free online trial of mmbert-safety-classifier-level1, you can try mmbert-safety-classifier-level1 online for free by clicking the link below.
llm-semantic-router mmbert-safety-classifier-level1 online free url in huggingface.co:
mmbert-safety-classifier-level1 is an open source model from GitHub that offers a free installation service, and any user can find mmbert-safety-classifier-level1 on GitHub to install. At the same time, huggingface.co provides the effect of mmbert-safety-classifier-level1 install, users can directly use mmbert-safety-classifier-level1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
mmbert-safety-classifier-level1 install url in huggingface.co: