PromptShield for Openclaw

A high-performance prompt injection firewall that protects AI agents from 113 threat patterns using multi-layered heuristic scoring.

stlas
v3.0.6
Feb 11, 2026
0
1.9k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install prompt-shield

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install prompt-shield using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is PromptShield?

PromptShield is a specialized security layer designed to safeguard AI agents against manipulative inputs and adversarial attacks. By utilizing a multi-layered pattern recognition system, it classifies incoming text into threat levels to prevent unauthorized command execution, memory poisoning, and identity overrides. As a developer-first tool within the Openclaw Skills ecosystem, it provides a zero-dependency solution for maintaining agent integrity.

Built by the RASSELBANDE collective, this firewall is battle-tested against real-world prompt injection scenarios. It serves as a critical defense mechanism for any developer building autonomous agents that interact with untrusted user input or external data sources, ensuring that the underlying LLM remains within its intended operational boundaries.

PromptShield Use Cases

  • Preventing prompt injection and jailbreak attempts in real-time agent interactions.
  • Sanitizing batch data or social media comments before they are processed by an AI model.
  • Protecting autonomous agents from memory poisoning and forced obedience attacks.
  • Implementing secure CLI hooks for tools like Claude Code to block malicious payloads.
  • Filtering out crypto spam, phishing links, and fake system authority messages from agent inputs.

How PromptShield Works

  1. The scanner accepts text input via CLI arguments, files, or standard input.
  2. The input is evaluated against 113 detection patterns spanning 14 distinct threat categories including command injection and social engineering.
  3. A heuristic engine calculates a danger score, applying bonus points if multiple attack vectors are detected simultaneously.
  4. The system checks a tamper-proof hash-chain whitelist to see if the specific pattern has been peer-reviewed and approved.
  5. Based on the final score (0-100), the tool returns a status of CLEAN, WARNING, or BLOCK to determine if the input should proceed.

PromptShield Setup

To get started with this component of your Openclaw Skills, ensure you have Python 3 installed and follow these steps:

# Install the only required dependency
pip install pyyaml

# Make the scanner executable
chmod +x shield.py

# Test a simple scan
./shield.py scan "Hello, how can I help you?"

For Claude Code integration, add the hook to your configuration:

{
  "hooks": {
    "UserInputSubmit": ["/path/to/prompt-shield/prompt-shield-hook.sh"]
  }
}

PromptShield Data Schema & Taxonomy

The skill manages its security logic through a structured YAML-based system:

File Description
patterns.yaml The primary database containing 113 patterns across 14 categories.
whitelist.yaml A hash-chained ledger of approved exceptions requiring peer signatures.
shield.py The main execution logic for scanning and heuristic scoring.

Threat classification is handled via the following scoring thresholds:

  • CLEAN (0-49): No significant threats detected; safe to process.
  • WARNING (50-79): Potential manipulation detected; proceed with caution.
  • BLOCK (80-100): High-confidence threat; input is rejected immediately.

PromptShield Advanced Features

  • Multi-layered heuristic combo detection that increases scores when diverse threat types (like fake authority and command injection) appear together.
  • Version 2 hash-chain whitelist which uses SHA256 integrity checks to prevent unauthorized local modifications to the security rules.
  • Mandatory peer review system for whitelisting, requiring at least two different agent IDs to approve any exception.
  • Support for multi-language detection, effectively identifying threats in English, German, Spanish, and French.
  • Batch processing capabilities with built-in duplicate detection for analyzing large datasets or logs. This makes it a versatile choice for those expanding their Openclaw Skills library.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*