AgentBench for OpenClaw for Openclaw

A comprehensive benchmarking suite designed to evaluate and quantify the real-world performance of AI agents using Openclaw Skills.

exe215
v1.0.0
Feb 22, 2026
1
1.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install agentbench

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install agentbench using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is AgentBench for OpenClaw?

AgentBench for OpenClaw is a professional-grade evaluation framework specifically built to measure the general capabilities of AI agents. Unlike standard coding benchmarks, it focuses on real-world operational efficiency across 40 tasks spanning 7 diverse domains. By utilizing this tool within your Openclaw Skills ecosystem, you can gain deep insights into how your agent handles file creation, complex research, data analysis, and multi-step workflows.

The skill provides a standardized way to test agent configurations, system prompts, and tool access. It ensures that your agent isn't just generating code, but is effectively navigating a workspace, managing errors, and adhering to complex instructions. Whether you are fine-tuning a local model or optimizing a cloud-based agent, AgentBench provides the data-driven validation necessary for high-performance Openclaw Skills.

AgentBench for OpenClaw Use Cases

  • Validating the reliability of new agent configurations or system prompt updates.
  • Comparing performance across different LLM backends within Openclaw Skills.
  • Identifying specific bottlenecks in tool usage, such as inefficient file reading or redundant command execution.
  • Generating standardized performance reports to track agent improvement over development cycles.
  • Auditing agent behavior for instruction adherence and error recovery capabilities.

How AgentBench for OpenClaw Works

  1. The skill scans the tasks directory to discover available benchmark tests based on difficulty levels and domain suites.
  2. A dedicated run directory is initialized with a unique timestamp to maintain an organized history of all performance data.
  3. The agent executes each assigned task in a sandboxed workspace, utilizing tools like web search, file editing, and command execution naturally.
  4. Real-time metrics are collected, including tool call counts, planning ratios, and execution duration for every task.
  5. A multi-layered scoring algorithm evaluates the results based on structural integrity, performance metrics, behavioral analysis, and final output quality.
  6. A comprehensive suite of reports, including an interactive HTML dashboard and an integrity-signed JSON file, is generated for analysis.

AgentBench for OpenClaw Setup

To get started with AgentBench, ensure your environment has the necessary system dependencies installed. This skill requires jq, bash, and python3 to be available in your path.

# Verify dependencies
jq --version
bash --version
python3 --version

Once dependencies are confirmed, you can list all available tasks to verify the installation of your Openclaw Skills suite:

/benchmark-list

AgentBench for OpenClaw Data Schema & Taxonomy

AgentBench organizes its output in the agentbench-results/{run-id}/ directory to ensure all Openclaw Skills data is traceable and structured.

File Format Purpose
results.json JSON Machine-readable summary including the integrity signature and overall scores.
report.html HTML A self-contained, interactive dashboard with color-coded performance visualizations.
report.md Markdown A human-readable summary of domain breakdowns and task failures.
scores.json JSON Detailed breakdown of the 4-layer scoring (Structural, Metrics, Behavioral, Quality).
metrics.json JSON Technical execution data including tool call frequency and timing.

AgentBench for OpenClaw Advanced Features

  • Support for strict mode execution to generate externally-verified scoring signatures for leaderboards.
  • Granular task filtering allowing users to run specific suites, individual tasks, or fast-track easy/medium tests.
  • Automated behavioral analysis that penalizes inefficient tool patterns, such as using cat instead of native read tools.
  • Side-by-side run comparison to visualize performance deltas between different agent versions.
  • Integrated workspace cleanup that preserves results while removing temporary execution artifacts.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Requires
Bins jqbashpython3
Github Stars: 0
forks: 0

Featured*