A comprehensive benchmarking suite designed to evaluate and quantify the real-world performance of AI agents using Openclaw Skills.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install agentbench
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install agentbench using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
AgentBench for OpenClaw is a professional-grade evaluation framework specifically built to measure the general capabilities of AI agents. Unlike standard coding benchmarks, it focuses on real-world operational efficiency across 40 tasks spanning 7 diverse domains. By utilizing this tool within your Openclaw Skills ecosystem, you can gain deep insights into how your agent handles file creation, complex research, data analysis, and multi-step workflows.
The skill provides a standardized way to test agent configurations, system prompts, and tool access. It ensures that your agent isn't just generating code, but is effectively navigating a workspace, managing errors, and adhering to complex instructions. Whether you are fine-tuning a local model or optimizing a cloud-based agent, AgentBench provides the data-driven validation necessary for high-performance Openclaw Skills.
To get started with AgentBench, ensure your environment has the necessary system dependencies installed. This skill requires jq, bash, and python3 to be available in your path.
# Verify dependencies
jq --version
bash --version
python3 --version
Once dependencies are confirmed, you can list all available tasks to verify the installation of your Openclaw Skills suite:
/benchmark-list
AgentBench organizes its output in the agentbench-results/{run-id}/ directory to ensure all Openclaw Skills data is traceable and structured.
| File | Format | Purpose |
|---|---|---|
results.json |
JSON | Machine-readable summary including the integrity signature and overall scores. |
report.html |
HTML | A self-contained, interactive dashboard with color-coded performance visualizations. |
report.md |
Markdown | A human-readable summary of domain breakdowns and task failures. |
scores.json |
JSON | Detailed breakdown of the 4-layer scoring (Structural, Metrics, Behavioral, Quality). |
metrics.json |
JSON | Technical execution data including tool call frequency and timing. |
Loading
A robust tool to capture fully rendered web pages and save them as structured blocks directly into Notion databases or pages.

A comprehensive AI creative studio skill that integrates 60+ models for video, image, and music generation through a single unified interface.

A smart LLM routing brain that automatically dispatches tasks to the most cost-effective and capable model across 20+ providers using a single API key.

A comprehensive AI video generation suite consolidating 37 industry-leading models into a single unified API for Openclaw Skills.

A professional integration for Moodle 4.x that enables AI agents to automate course management, enrollments, and grading via REST APIs.

A specialized social networking protocol for AI agents to interact, share, and discover content through a structured API.








































