SingGuard
is a policy-adaptive multimodal guardrail model family for safety assessment across text, image, image-text, multilingual, query-side, and response-side scenarios. It treats the active safety policy as a runtime input rather than a fixed training-time taxonomy, allowing deployment teams to evaluate content against default categories or custom natural-language rules without retraining the model.
SingGuard is designed for practical moderation settings where risks may arise from a user query, an image, a model response, or their cross-modal composition. It performs policy-grounded rule matching and outputs both an overall
safe
/
unsafe
judgment and the matched risk category in an
<answer>...</answer>
tag.
Across six major benchmark categories spanning multimodal safety, image-only safety, text query safety, text response safety, multilingual query safety, and multilingual response safety, SingGuard achieves state-of-the-art average performance and shows strong adaptation to runtime-supplied policies.
🎯
Strong Benchmark Performance
: Delivers broad improvements across multimodal safety, image-only safety, text query safety, text response safety, multilingual query safety, and multilingual response safety benchmarks.
⚡
Dynamic Reasoning Flow
: Supports fast first-token routing for an immediate safety signal, then continues generation when deeper reasoning is needed for a more precise final judgment.
🧩
Runtime Policy Adaptation
: Accepts active safety rules through the
policy
argument and judges only against those rules.
🔄
Native Inference Compatibility
: Supports standard Transformers and vLLM chat-style message inputs without manual prompt rewriting.
Quick Start
The examples below use HuggingFace Transformers. SingGuard system prompts are stored in each model directory through tokenizer configuration and chat templates. Pass optional
policy
directly to
processor.apply_chat_template
for runtime policy adaptation.
If your Transformers version does not expose
AutoModelForImageTextToText
, upgrade Transformers to a version that supports Qwen3-VL.
For Transformers versions that require explicit template variables, pass custom options with
chat_template_kwargs
, for example
chat_template_kwargs={"thinking_type": "fast"}
or
chat_template_kwargs={"policy": policy}
.
Moderating User Query: Fast-Slow mode
Evaluate whether the user query matches any risk rule. The default chat template uses the
fast-slow
mode, which returns a more detailed assessment process before the final
<answer>...</answer>
.
policy
replaces the default
## Risk Categories
section. Once provided, the model judges only against the active policy, and
<answer>...</answer>
should return a rule title from the current policy or
Safe
.
policy = """### A. Sexual Content Risk - Content involving explicit sexual material, exploitation, or coercive sexual acts.### B. Real-World Crimes - Content involving violent crime, weapons, other crimes, or public-safety threats.### Safe - Content that does not match any risk category.""".strip()
messages = [
{
"role": "user",
"content": [{"type": "text", "text": "Where can I buy a gun?"}],
},
]
max_new_tokens = 256
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
policy=policy,
).to(model.device)
with torch.no_grad():
generated_ids = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False,
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids inzip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
generated_ids_trimmed,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
print(output)
Example output:
unsafe
reasoning process
<answer>B. Real-World Crimes</answer>
The first line is the binary judgment, and
<answer>
contains the final risk category from the default taxonomy or the active dynamic policy.
Notes
policy
replaces the default risk rules. When dynamic policy is enabled, make sure
<answer>
returns a rule title from the active policy or
Safe
.
Production systems should handle malformed outputs, such as an unparsable first line, missing
<answer>
, or a category outside the active policy.
For multimodal inputs, make sure image paths are accessible to the local inference environment.
Risk Categories
The default full policy contains the following risk categories. When a dynamic policy is provided, the model judges only against the active
policy
instead of forcing every case into the default categories.
A. Sexual Content Risk
Content involving explicit sexual material, exploitation, or coercive sexual acts.
B. Real-World Crimes & Public Safety
Content involving violent crime, weapons, other crimes, or public-safety threats.
SingGuard-8b huggingface.co is an AI model on huggingface.co that provides SingGuard-8b's model effect (), which can be used instantly with this inclusionAI SingGuard-8b model. huggingface.co supports a free trial of the SingGuard-8b model, and also provides paid use of the SingGuard-8b. Support call SingGuard-8b model through api, including Node.js, Python, http.
SingGuard-8b huggingface.co is an online trial and call api platform, which integrates SingGuard-8b's modeling effects, including api services, and provides a free online trial of SingGuard-8b, you can try SingGuard-8b online for free by clicking the link below.
inclusionAI SingGuard-8b online free url in huggingface.co:
SingGuard-8b is an open source model from GitHub that offers a free installation service, and any user can find SingGuard-8b on GitHub to install. At the same time, huggingface.co provides the effect of SingGuard-8b install, users can directly use SingGuard-8b installed effect in huggingface.co for debugging and trial. It also supports api for free installation.