A performance-tracking ecosystem that benchmarks AI agent skills to drive continuous self-improvement and reliability.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install skillbench
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install skillbench using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
skillbench provides a robust framework for monitoring the lifecycle of capabilities within the Openclaw Skills ecosystem. It serves as a dedicated benchmarking layer that allows developers to track versioned performance metrics, compare improvements between releases, and identify specific areas where an agent's skills may be declining or failing. By formalizing the feedback loop between task execution and performance analysis, it ensures that AI agents evolve based on data rather than assumptions.
As part of a wider technical stack including ClawVault and tasktime, skillbench enables deep insights into agent behavior. It translates raw execution data into actionable grades, helping developers maintain high standards for their Openclaw Skills. Whether you are debugging a complex workflow or optimizing a production agent, this tool provides the metrics necessary to validate that every code change results in a measurable improvement.
To get started with this tool for managing Openclaw Skills, install the CLI package via npm:
npm install -g @versatly/skillbench
Once installed, you can synchronize your existing environment:
skillbench sync --all
The skill organizes its benchmarking data using the following taxonomy to ensure high-quality reporting for Openclaw Skills:
| Component | Details |
|---|---|
| Skill Identity | Name, version (semver), and source metadata. |
| Performance Metrics | Success Rate (40%), Avg Duration (30%), Consistency (20%), and Trend (10%). |
| Grading Schema | Letter grades from A+ (95-100) down to D (<50). |
| Status Signals | Flags such as 'needs work', 'stale', or 'declining' based on recent activity. |
| Export Formats | Support for Markdown, JSON, and HTML dashboard outputs. |
Loading
pdauth is a dynamic OAuth bridge that allows AI agents to request authorization and call tools across 2500+ external APIs via Pipedream.

A powerful command-line interface for generating, editing, and managing images using OpenAI's state-of-the-art GPT Image and DALL-E models.

A powerful automation skill for managing LinkedIn profiles and organization pages through Pipedream OAuth integration.

A comprehensive command-line interface for managing Clover POS operations, from inventory tracking to advanced financial reporting.

A high-performance CLI task timer and benchmarking tool designed for AI agents to track learning progression and automate productivity logging.

A social evolution experiment where AI agents compete for survival through a unique citation-based energy economy.








































