skillbench for Openclaw

A performance-tracking ecosystem that benchmarks AI agent skills to drive continuous self-improvement and reliability.

g9pedro
v2.0.0
Feb 11, 2026
0
2.4k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install skillbench

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install skillbench using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is skillbench?

skillbench provides a robust framework for monitoring the lifecycle of capabilities within the Openclaw Skills ecosystem. It serves as a dedicated benchmarking layer that allows developers to track versioned performance metrics, compare improvements between releases, and identify specific areas where an agent's skills may be declining or failing. By formalizing the feedback loop between task execution and performance analysis, it ensures that AI agents evolve based on data rather than assumptions.

As part of a wider technical stack including ClawVault and tasktime, skillbench enables deep insights into agent behavior. It translates raw execution data into actionable grades, helping developers maintain high standards for their Openclaw Skills. Whether you are debugging a complex workflow or optimizing a production agent, this tool provides the metrics necessary to validate that every code change results in a measurable improvement.

skillbench Use Cases

  • Benchmarking different versions of a skill to validate performance gains before deployment.
  • Identifying declining success rates in Openclaw Skills caused by API shifts or prompt drift.
  • Automatically capturing task duration and success metrics via integration with tasktime.
  • Comparing the performance of multiple AI agents side-by-side using a leaderboard.
  • Establishing performance baselines to prevent regressions in CI/CD pipelines.

How skillbench Works

  1. Initialize a specific skill and version to track using the CLI to set the active context.
  2. Execute a specific task while the skill is active, utilizing integrated timers for accuracy.
  3. Record the result of the operation, including success/failure status and duration, which is then stored as a benchmark.
  4. Analyze the collected data using the grading system to see scores based on success rates and consistency.
  5. Review improvement signals to determine if the skill requires maintenance or version bumping.
  6. Sync benchmarks to ClawVault to maintain a persistent performance history across sessions.

skillbench Setup

To get started with this tool for managing Openclaw Skills, install the CLI package via npm:

npm install -g @versatly/skillbench

Once installed, you can synchronize your existing environment:

skillbench sync --all

skillbench Data Schema & Taxonomy

The skill organizes its benchmarking data using the following taxonomy to ensure high-quality reporting for Openclaw Skills:

Component Details
Skill Identity Name, version (semver), and source metadata.
Performance Metrics Success Rate (40%), Avg Duration (30%), Consistency (20%), and Trend (10%).
Grading Schema Letter grades from A+ (95-100) down to D (<50).
Status Signals Flags such as 'needs work', 'stale', or 'declining' based on recent activity.
Export Formats Support for Markdown, JSON, and HTML dashboard outputs.

skillbench Advanced Features

  • CI/CD Integration: Use automated baseline checks that exit with error codes upon performance regression.
  • Regression Detection: Set fixed baselines to ensure that updates to Openclaw Skills never compromise existing efficiency.
  • Automated Testing: Run scheduled smoke tests and full suites to monitor agent health continuously.
  • HTML Dashboards: Generate visual performance reports and open them directly in the browser for stakeholders.
  • Multi-Agent Leaderboards: Compare various agent configurations to determine which handles specific Openclaw Skills most effectively.

SKILL.md


Loading

Related Openclaw Skills

Featured*