Content Moderation Skill for Openclaw

A two-layer safety framework designed to protect AI agents from prompt injections and sensitive content violations.

zskyx
v1.0.0
Feb 2, 2026
5
3k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install detect-injection

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install detect-injection using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Content Moderation Skill?

The Content Moderation skill provides a critical security infrastructure for autonomous agents. It implements a dual-layer defense mechanism that inspects both incoming user messages and outgoing agent responses. By leveraging high-confidence classifiers and specialized moderation endpoints, it ensures that agents remain within safety boundaries and resist adversarial attempts to manipulate their internal logic. Integrating this capability into your Openclaw Skills allows for the safe deployment of LLMs in public, multi-user, or enterprise environments where data integrity and content policy compliance are paramount.

This skill is particularly effective at identifying sophisticated prompt injection attacks that attempt to bypass system instructions. It uses the ProtectAI DeBERTa classifier to provide binary safe/injection verdicts with extremely high confidence, while optionally utilizing standard moderation APIs to filter for 13 distinct categories of prohibited content, including harassment and hate speech.

Content Moderation Skill Use Cases

  • Detecting and blocking prompt injection attempts that try to override agent instructions.
  • Preventing the leakage of internal configurations or hidden system prompts.
  • Moderating messages from untrusted users in public or group chat environments.
  • Ensuring agent-generated responses do not contain violence, self-harm, or sexual content.
  • Automated rewriting of flagged output to comply with safety policies.

How Content Moderation Skill Works

  1. The agent captures the text intended for processing or the response it has just generated.
  2. The system invokes the scripts/moderate.sh utility with the appropriate direction flag (input or output).
  3. For user inputs, a ProtectAI DeBERTa model via HuggingFace Inference performs a prompt injection check.
  4. If configured, the text is sent to the OpenAI omni-moderation endpoint to evaluate against safety categories.
  5. The script generates a JSON report containing confidence scores and a boolean flagged status.
  6. The agent interprets the JSON result to either proceed, decline the request, or sanitize its own output.

Content Moderation Skill Setup

To enable this skill within your Openclaw Skills environment, you must export the necessary API tokens and configuration variables:

export HF_TOKEN="hf_..."           # Required: Get from huggingface.co/settings/tokens
export OPENAI_API_KEY="sk-..."     # Optional: Enables the content safety layer
export INJECTION_THRESHOLD="0.85"  # Optional: Adjust sensitivity (default is 0.85)

Content Moderation Skill Data Schema & Taxonomy

The moderation script returns a structured JSON object with the following metadata taxonomy:

Key Type Description
flagged Boolean Final verdict; true if any safety layer triggers.
direction String Context of the check: either input or output.
injection Object Includes flagged (bool) and score (float) for injection detection.
content Object Detailed category flagging from the moderation API.
action String A suggested instruction for the agent on how to handle the failure.

Content Moderation Skill Advanced Features

  • Multi-layer defense combining dedicated injection classifiers with broad content moderation.
  • High-confidence detection (99.99%) for typical prompt injection attacks using DeBERTa.
  • Flexible threshold management to balance between security and false positives.
  • Free-tier compatibility using HuggingFace and OpenAI's moderation endpoints.
  • Automated action strings to guide agent behavior during security events.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*