PwnClaw Security Scan for Openclaw

A comprehensive security testing suite designed to harden AI agents against prompt injection, jailbreaks, and advanced adversarial attacks.

gemini2027
v1.0.0
Feb 10, 2026
2
2.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install pwnclaw-security-scan

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install pwnclaw-security-scan using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is PwnClaw Security Scan?

PwnClaw Security Scan is a specialized diagnostic skill that subjects AI agents to over 112 real-world attacks across 14 distinct categories. It is designed to identify critical vulnerabilities like indirect injection, agency hijacking, and tool poisoning, which are common risks in modern agentic workflows. By integrating this skill into your development lifecycle, you ensure your agents remain resilient against sophisticated exploits.

As part of the Openclaw Skills ecosystem, PwnClaw provides actionable intelligence through a security score and tailored fix instructions. This allows developers to iteratively harden their agents system prompts and logic based on empirical data rather than guesswork, ensuring a secure and reliable user experience.

PwnClaw Security Scan Use Cases

  • Hardening system prompts against adversarial jailbreak attempts before production deployment.
  • Periodic auditing of multi-agent systems to prevent cross-agent data exfiltration or memory poisoning.
  • Verifying the security of Model Context Protocol (MCP) tool integrations against malicious poisoning.
  • Assessing agent susceptibility to social engineering and sycophancy in open-ended user interactions.

How PwnClaw Security Scan Works

  1. The user initiates a scan via the PwnClaw dashboard or API using the Openclaw Skills framework.
  2. The system delivers a series of targeted adversarial prompts to the agent, ranging from simple injections to complex obfuscation techniques.
  3. The agent processes these prompts and returns responses, which are then analyzed for compliance or failure signatures.
  4. PwnClaw calculates a security score and generates specific mitigation rules.
  5. The developer applies the recommended security instructions to the agent's system prompt and re-tests to confirm the vulnerability is closed.

PwnClaw Security Scan Setup

To begin using PwnClaw Security Scan, follow these steps:

  1. Sign up for a free account at https://www.pwnclaw.com.
  2. If using the Manual API mode, retrieve your test token from the dashboard.
  3. Implement the following logic for automated testing via CLI or script:
# Fetch the next attack prompt
curl -X GET "https://www.pwnclaw.com/api/test/{token}"

# Submit the agent's response for analysis
curl -X POST "https://www.pwnclaw.com/api/test/{token}" \
     -H "Content-Type: application/json" \
     -d '{"response": "YOUR_AGENT_RESPONSE"}'
  1. For agents with public HTTP endpoints, simply provide the URL in the PwnClaw dashboard for fully automatic scanning.

PwnClaw Security Scan Data Schema & Taxonomy

The skill generates structured reports based on the following taxonomy:

Category Description
Injection Tests for prompt and indirect injection vulnerabilities.
Jailbreaks Evaluates refusal bypass and safety filter triggers.
Agency Checks for agency hijacking and unauthorized tool usage.
MCP/Tool Specifically tests for poisoning of Model Context Protocol integrations.
Scoring A numerical representation of the agent's current security posture.

All results, including the transcript of the attack and the generated fix instructions, are stored and accessible via the PwnClaw dashboard.

PwnClaw Security Scan Advanced Features

  • Multi-Turn Attack Sequences: Simulates complex, stateful interactions to find deep-seated logic flaws.
  • Custom Security Rule Generation: Automatically creates hardened system prompt snippets based on failed test cases.
  • API-First Integration: Easily integrate security checks into CI/CD pipelines via RESTful endpoints.
  • Publicly Auditable Core: Open-source foundation ensuring transparency in the testing methodology and alignment with Openclaw Skills standards.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*