llm-semantic-router / mmbert-safety-classifier-level1

huggingface.co
Total runs: 1
24-hour runs: 0
7-day runs: -1
30-day runs: -10
Model's Last Updated: January 22 2026
text-classification

Introduction of mmbert-safety-classifier-level1

Model Details of mmbert-safety-classifier-level1

mmBERT Safety Classifier - Level 1 (Binary)

A binary safety classifier for detecting unsafe content in LLM inputs. Part of a hierarchical MLCommons-aligned safety classification system.

Model Description

This is Level 1 of a 2-level hierarchical safety classifier:

  • Level 1 (this model) : Binary classification (safe/unsafe) - high recall for catching threats
  • Level 2 : 9-class hazard taxonomy (MLCommons AI Safety aligned) - for categorizing unsafe content
Performance
Metric Score
Accuracy 84.9%
F1 Score 84.9%
Labels
ID Label Description
0 safe Content is safe
1 unsafe Content is potentially harmful
Training Hyperparameters
Parameter Value
Base Model jhu-clsp/mmBERT-base
LoRA Rank 32
LoRA Alpha 64
LoRA Dropout 0.1
Learning Rate 5e-5
Epochs 10
Batch Size 64
Max Samples 18,000
Training Samples 12,600
Validation Samples 2,700
Usage
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel

# Load model
tokenizer = AutoTokenizer.from_pretrained("llm-semantic-router/mmbert-safety-classifier-level1")
base_model = AutoModelForSequenceClassification.from_pretrained(
    "jhu-clsp/mmBERT-base", 
    num_labels=2,
    torch_dtype=torch.float32
)
model = PeftModel.from_pretrained(base_model, "llm-semantic-router/mmbert-safety-classifier-level1")
model.eval()

# Inference
text = "How do I bake a chocolate cake?"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)

with torch.no_grad():
    outputs = model(**inputs)
    pred = outputs.logits.argmax(-1).item()
    
labels = {0: "safe", 1: "unsafe"}
print(f"Prediction: {labels[pred]}")
Hierarchical Classification Pipeline

For complete safety classification, use with Level 2:

# If Level 1 predicts "unsafe", run Level 2 for hazard category
if pred == 1:  # unsafe
    # Load and run Level 2 model for specific hazard category
    # See: llm-semantic-router/mmbert-safety-classifier-level2
    pass
Training Data
  • Primary : nvidia/Aegis-AI-Content-Safety-Dataset-2.0
  • Enhanced : 110 edge case examples for underrepresented categories (CSE, medical advice, misinformation)
  • Balance : 50/50 safe/unsafe split with oversampling
Intended Use
  • Content moderation for LLM applications
  • Input filtering for AI safety systems
  • Guardrail implementation for chatbots and AI assistants
Limitations
  • May miss subtle misinformation content
  • Trained primarily on English content
  • Should be used as part of a defense-in-depth strategy
Citation
@misc{mmbert-safety-classifier,
  title={mmBERT Safety Classifier},
  author={LLM Semantic Router Team},
  year={2026},
  publisher={HuggingFace},
  url={https://huggingface.co/llm-semantic-router/mmbert-safety-classifier-level1}
}

Runs of llm-semantic-router mmbert-safety-classifier-level1 on huggingface.co

1
Total runs
0
24-hour runs
0
3-day runs
-1
7-day runs
-10
30-day runs

More Information About mmbert-safety-classifier-level1 huggingface.co Model

More mmbert-safety-classifier-level1 license Visit here:

https://choosealicense.com/licenses/apache-2.0

mmbert-safety-classifier-level1 huggingface.co

mmbert-safety-classifier-level1 huggingface.co is an AI model on huggingface.co that provides mmbert-safety-classifier-level1's model effect (), which can be used instantly with this llm-semantic-router mmbert-safety-classifier-level1 model. huggingface.co supports a free trial of the mmbert-safety-classifier-level1 model, and also provides paid use of the mmbert-safety-classifier-level1. Support call mmbert-safety-classifier-level1 model through api, including Node.js, Python, http.

llm-semantic-router mmbert-safety-classifier-level1 online free

mmbert-safety-classifier-level1 huggingface.co is an online trial and call api platform, which integrates mmbert-safety-classifier-level1's modeling effects, including api services, and provides a free online trial of mmbert-safety-classifier-level1, you can try mmbert-safety-classifier-level1 online for free by clicking the link below.

llm-semantic-router mmbert-safety-classifier-level1 online free url in huggingface.co:

https://huggingface.co/llm-semantic-router/mmbert-safety-classifier-level1

mmbert-safety-classifier-level1 install

mmbert-safety-classifier-level1 is an open source model from GitHub that offers a free installation service, and any user can find mmbert-safety-classifier-level1 on GitHub to install. At the same time, huggingface.co provides the effect of mmbert-safety-classifier-level1 install, users can directly use mmbert-safety-classifier-level1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

mmbert-safety-classifier-level1 install url in huggingface.co:

https://huggingface.co/llm-semantic-router/mmbert-safety-classifier-level1

Url of mmbert-safety-classifier-level1

mmbert-safety-classifier-level1 huggingface.co Url

Provider of mmbert-safety-classifier-level1 huggingface.co

llm-semantic-router
ORGANIZATIONS

Other API from llm-semantic-router