A comprehensive auditing framework that combines automated structural checks with a 25-point manual rubric to ensure AI agent skills are production-ready.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install skill-evaluator
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install skill-evaluator using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The Skill Evaluator is a professional-grade assessment tool designed to validate the quality, reliability, and security of agent capabilities. It utilizes a hybrid approach, pairing automated Python-based structural analysis with a deep manual rubric derived from industry standards like ISO 25010, OpenSSF, and Shneiderman heuristics. This ensures that Openclaw Skills meet rigorous requirements for functional suitability and performance before they are deployed or published.
By providing a standardized scoring system and templated reporting, the evaluator helps developers identify critical blockers (P0) and necessary improvements (P1). It acts as a quality gate, ensuring that every integration is maintainable, secure against credential leaks, and provides a seamless experience for both AI agents and human users.
To start evaluating Openclaw Skills, ensure you have Python 3.6+ and the required YAML parser installed:
pip install pyyaml
You can then run the evaluator script against any skill directory using the following commands:
# Basic automated check
python3 scripts/eval-skill.py /path/to/skill
# Detailed check with verbose logging
python3 scripts/eval-skill.py /path/to/skill --verbose
# Export results in machine-readable format
python3 scripts/eval-skill.py /path/to/skill --json
The evaluator generates structured quality data that maps to specific reliability frameworks. The findings are typically organized as follows:
| Component | Type | Description |
|---|---|---|
| Automated Score | Numeric | Derived from file structure, syntax, and dependency audits. |
| Manual Rubric | Table | 25 criteria scored 0–4 based on the reference rubric. |
| Severity Matrix | P0-P2 | Priority-based classification of bugs and design flaws. |
| Verdict | String | Final status based on score (e.g., Excellent, Acceptable, Not Ready). |
| EVAL.md | File | The final markdown report generated within the skill directory. |
Loading
An AI-powered document automation tool for Feishu and Lark that handles long-form content, smart categorization, and ownership transfer seamlessly.

A comprehensive toolset for automated prediction market trading on the Binance Smart Chain using the 0xProbable CLOB protocol.

A high-performance trading toolkit for the 0xProbable prediction market on the BSC mainnet, enabling automated event searching and order book execution.

A sophisticated message history analyzer that extracts relationship dynamics and communication patterns from iMessage and Signal data.

A comprehensive automation tool for managing Zotero reference libraries, fetching metadata, and organizing academic citations via the Web API.

A professional-grade Binance trading skill equipped with AI market analysis and automated risk management.








































