A two-layer safety framework designed to protect AI agents from prompt injections and sensitive content violations.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install detect-injection
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install detect-injection using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The Content Moderation skill provides a critical security infrastructure for autonomous agents. It implements a dual-layer defense mechanism that inspects both incoming user messages and outgoing agent responses. By leveraging high-confidence classifiers and specialized moderation endpoints, it ensures that agents remain within safety boundaries and resist adversarial attempts to manipulate their internal logic. Integrating this capability into your Openclaw Skills allows for the safe deployment of LLMs in public, multi-user, or enterprise environments where data integrity and content policy compliance are paramount.
This skill is particularly effective at identifying sophisticated prompt injection attacks that attempt to bypass system instructions. It uses the ProtectAI DeBERTa classifier to provide binary safe/injection verdicts with extremely high confidence, while optionally utilizing standard moderation APIs to filter for 13 distinct categories of prohibited content, including harassment and hate speech.
scripts/moderate.sh utility with the appropriate direction flag (input or output).To enable this skill within your Openclaw Skills environment, you must export the necessary API tokens and configuration variables:
export HF_TOKEN="hf_..." # Required: Get from huggingface.co/settings/tokens
export OPENAI_API_KEY="sk-..." # Optional: Enables the content safety layer
export INJECTION_THRESHOLD="0.85" # Optional: Adjust sensitivity (default is 0.85)
The moderation script returns a structured JSON object with the following metadata taxonomy:
| Key | Type | Description |
|---|---|---|
flagged |
Boolean | Final verdict; true if any safety layer triggers. |
direction |
String | Context of the check: either input or output. |
injection |
Object | Includes flagged (bool) and score (float) for injection detection. |
content |
Object | Detailed category flagging from the moderation API. |
action |
String | A suggested instruction for the agent on how to handle the failure. |
Loading
An intelligent multi-model routing manager that optimizes AI performance and reduces costs through automated model selection and failover logic.

RTFM Testing is a methodology that spawns fresh AI agents with zero context to validate whether documentation is actually usable for its intended tasks.

Bagman is a comprehensive security framework for AI agents to manage private keys and API secrets without risking accidental exposure or theft.

Bagman is a secure key management framework for AI agents to handle wallets, API secrets, and private keys safely.

A smart PDF utility that automatically detects page orientation to apply perfectly centered, adjustable text watermarks.

A comprehensive developer tool for building .NET applications that process Excel, Word, PDF, and email files using GemBox components.








































