Salamandra-7b-instruct-guard is a state-of-the-art safety classification model designed for Catalan, Spanish, and English content moderation. Built on the Salamandra-7B-Instruct foundation model from Barcelona Supercomputing Center (BSC), it provides both binary (safe/unsafe) and multiclass safety classification capabilities.
Model Details
Model Description
Salamandra-7b-instruct-guard is a fine-tuned safety guardrail model that classifies text content across multiple safety categories. The model addresses three primary categories of unsafe content: Dangerous Content, Toxic Content, and Sexual Content
The model is specifically optimized for European languages, with particular emphasis on Catalan—a language often overlooked in existing safety models—alongside Spanish and English.
Developed by:
Barcelona Supercomputing Center (BSC)
Model type:
Text Classification (Safety Guardrail)
LoRA (Low-Rank Adaptation) applied to all attention layers
Maximum sequence length: 8,192 tokens
Training objectives: Binary classification and multiclass classification
Two separate models trained: Binary and Multiclass variants
Training hyperparameters:
Optimizer:
AdamW
Learning rate:
5e-5
Training epochs:
3
Batch processing:
Distributed Data Parallel across 4× A100 GPUs
Framework:
Hugging Face Accelerate
Loss function:
Cross-entropy (single-label) / CausalLM (generative classification)
Evaluation
Testing Data
Test set:
722 samples from SalGuard_10k
Catalan: 210 samples
Spanish: 211 samples
English: 300 samples
Metrics
Binary Classification:
Accuracy:
0.798
Weighted F1:
0.797
Multiclass Classification (C0-C3):
Accuracy:
0.733
Weighted F1:
0.731
Per-class F1 scores (Multiclass):
C0 (Safe): 0.75
C1 (Dangerous): 0.75
C2 (Toxic): 0.64
C3 (Sexual): 0.78
Results
Comparison with Baseline Models
Binary Classification (Weighted F1):
Model
Weighted F1
Salamandra-7b-instruct-guard (Binary)
0.797
LlamaGuard 3
0.717
Salamandra 7B (base)
0.661
ShieldGemma 9B
0.638
Multiclass Classification (Weighted F1):
Model
Weighted F1
Salamandra-7b-instruct-guard (Multiclass)
0.731
LlamaGuard 3
0.618
ShieldGemma 9B
0.607
SalGuard V1
0.568
Salamandra 7B (base)
0.420
Key improvements:
+21% relative improvement
over base model in binary classification
+76% relative improvement
over base model in multiclass accuracy
+28% relative improvement
over SalGuard V1 in multiclass classification
Language-Specific Performance
The model demonstrates robust cross-language performance:
Binary (Weighted F1):
Catalan: 0.796
Spanish: 0.798
Multiclass (Weighted F1):
Catalan: 0.728
Spanish: 0.734
Bias, Risks, and Limitations
Known Limitations
Adversarial robustness:
The model focuses on moderating LLM responses and is not robust against adversarial unsafe user requests
Toxic Content (C2) detection:
Lower performance on hate speech, harassment, and profanity compared to dangerous and sexual content categories
False positives:
May flag safe discussions about dangerous topics (e.g., news about crimes) as harmful content
Context sensitivity:
Struggles with subtle contextual distinctions between discussing harmful topics and promoting them
Crowdsourced data quality:
Subset of training data proofread by crowdworkers may have variable linguistic quality compared to expert-reviewed samples
Annotation disagreement:
Significant disagreement exists between human annotators and between humans and LLM judges, reflecting inherent subjectivity in safety classification
Bias Considerations
Cultural adaptation:
Profanity (S6) definitions are culturally adapted for Catalan and Spanish contexts
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project ILENIA with reference 2022/TL22/00215337.
Runs of BSC-LT salamandra-7b-instruct-guard on huggingface.co
103
Total runs
0
24-hour runs
-1
3-day runs
-3
7-day runs
47
30-day runs
More Information About salamandra-7b-instruct-guard huggingface.co Model
More salamandra-7b-instruct-guard license Visit here:
salamandra-7b-instruct-guard huggingface.co is an AI model on huggingface.co that provides salamandra-7b-instruct-guard's model effect (), which can be used instantly with this BSC-LT salamandra-7b-instruct-guard model. huggingface.co supports a free trial of the salamandra-7b-instruct-guard model, and also provides paid use of the salamandra-7b-instruct-guard. Support call salamandra-7b-instruct-guard model through api, including Node.js, Python, http.
salamandra-7b-instruct-guard huggingface.co is an online trial and call api platform, which integrates salamandra-7b-instruct-guard's modeling effects, including api services, and provides a free online trial of salamandra-7b-instruct-guard, you can try salamandra-7b-instruct-guard online for free by clicking the link below.
BSC-LT salamandra-7b-instruct-guard online free url in huggingface.co:
salamandra-7b-instruct-guard is an open source model from GitHub that offers a free installation service, and any user can find salamandra-7b-instruct-guard on GitHub to install. At the same time, huggingface.co provides the effect of salamandra-7b-instruct-guard install, users can directly use salamandra-7b-instruct-guard installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
salamandra-7b-instruct-guard install url in huggingface.co: