Skill Evaluator for Openclaw

A comprehensive auditing framework that combines automated structural checks with a 25-point manual rubric to ensure AI agent skills are production-ready.

terwox
v1.0.0
Jan 31, 2026
3
3.3k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install skill-evaluator

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install skill-evaluator using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Skill Evaluator?

The Skill Evaluator is a professional-grade assessment tool designed to validate the quality, reliability, and security of agent capabilities. It utilizes a hybrid approach, pairing automated Python-based structural analysis with a deep manual rubric derived from industry standards like ISO 25010, OpenSSF, and Shneiderman heuristics. This ensures that Openclaw Skills meet rigorous requirements for functional suitability and performance before they are deployed or published.

By providing a standardized scoring system and templated reporting, the evaluator helps developers identify critical blockers (P0) and necessary improvements (P1). It acts as a quality gate, ensuring that every integration is maintainable, secure against credential leaks, and provides a seamless experience for both AI agents and human users.

Skill Evaluator Use Cases

  • Performing pre-publish audits for Openclaw Skills to ensure they meet security and reliability benchmarks.
  • Benchmarking third-party agent skills against the ISO 25010 framework for functional correctness.
  • Standardizing code reviews for internal developer teams building complex AI workflows.
  • Identifying security vulnerabilities and credential risks in skill scripts before deployment.

How Skill Evaluator Works

  1. Run the automated eval-skill.py script to scan for structural integrity, syntax errors, and documented environment variables.
  2. Review the core logic and scripts of the target skill to assess code quality and error-handling capabilities.
  3. Apply the 25-criteria rubric to assign scores (0–4) across eight distinct categories, including Security and Agent-Specific heuristics.
  4. Prioritize findings based on severity, categorizing them into P0 (blocking), P1 (recommended), or P2 (optional) fixes.
  5. Generate a standardized EVAL.md report using the provided template to document the final score and verdict.

Skill Evaluator Setup

To start evaluating Openclaw Skills, ensure you have Python 3.6+ and the required YAML parser installed:

pip install pyyaml

You can then run the evaluator script against any skill directory using the following commands:

# Basic automated check
python3 scripts/eval-skill.py /path/to/skill

# Detailed check with verbose logging
python3 scripts/eval-skill.py /path/to/skill --verbose

# Export results in machine-readable format
python3 scripts/eval-skill.py /path/to/skill --json

Skill Evaluator Data Schema & Taxonomy

The evaluator generates structured quality data that maps to specific reliability frameworks. The findings are typically organized as follows:

Component Type Description
Automated Score Numeric Derived from file structure, syntax, and dependency audits.
Manual Rubric Table 25 criteria scored 0–4 based on the reference rubric.
Severity Matrix P0-P2 Priority-based classification of bugs and design flaws.
Verdict String Final status based on score (e.g., Excellent, Acceptable, Not Ready).
EVAL.md File The final markdown report generated within the skill directory.

Skill Evaluator Advanced Features

  • Multi-framework compliance checking including ISO 25010, OpenSSF, and Shneiderman heuristics.
  • Automated credential and environment variable scanning to prevent sensitive data leaks.
  • Progressive disclosure and idempotency checks specifically tailored for AI agent behaviors.
  • Integration support for SkillLens to perform deep security scans for prompt injection and privilege bypass.
  • JSON output mode for integrating skill quality checks into CI/CD pipelines.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*