HLE Benchmark Evolver for Openclaw

Automates HLE benchmark reward ingestion and curriculum generation for AI capability evolution.

wanng-ide
v1.0.0
Feb 16, 2026
1
1.8k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install hle-benchmark-evolver

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install hle-benchmark-evolver using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is HLE Benchmark Evolver?

The HLE Benchmark Evolver is a specialized utility designed to operationalize Humanity's Last Exam (HLE) score-driven evolution. As part of the ecosystem of Openclaw Skills, it bridges the gap between raw benchmark results and actionable agent improvements. It allows developers to ingest question-level output and transform it into high-fidelity reward signals and curriculum queues, ensuring the agent focuses on the most impactful learning paths first.

This skill is particularly effective for teams looking to systematically increase their benchmark scores through automated feedback loops. By integrating with the capability-evolver, it provides a structured approach to model refinement, prioritizing easy-first curriculum stages and identifying specific focus subjects or modalities that require attention.

HLE Benchmark Evolver Use Cases

  • Improving Humanity's Last Exam (HLE) scores through targeted evolution.
  • Converting raw JSON benchmark reports into granular reward data.
  • Implementing an easy-first curriculum to stabilize agent learning.
  • Generating real-time benchmark progress snapshots and trend analysis.
  • Automating multi-cycle evolution loops with external evaluators.

How HLE Benchmark Evolver Works

  1. Validate the provided HLE report JSON to ensure schema compliance and data integrity.
  2. Ingest the report data into the capability-evolver benchmark reward state engine.
  3. Generate specific curriculum signals including focus subjects, modalities, and question-level priorities.
  4. Produce a compact JSON summary representing the current run accuracy and reward trends.
  5. Trigger subsequent evolution or solidification cycles based on the generated curriculum signals.

HLE Benchmark Evolver Setup

To get started with this component of Openclaw Skills, ensure you have your HLE benchmark reports ready. Run the basic ingestion using the following command:

node skills/hle-benchmark-evolver/run_result.js --report=/absolute/path/hle_report.json

For a full automated evolution cycle, use the pipeline script:

node skills/hle-benchmark-evolver/run_pipeline.js --report=/absolute/path/hle_report.json --cycles=1

You can also let the pipeline generate reports by providing an evaluation command:

node skills/hle-benchmark-evolver/run_pipeline.js --report=/path/hle_report.json --eval_cmd="python /path/to/eval_hle.py --out {{report}}" --cycles=3

HLE Benchmark Evolver Data Schema & Taxonomy

The skill outputs a structured JSON contract containing the following metrics to track evolution progress:

Field Description
benchmark_id Identifier for the benchmark (e.g., cais/hle)
accuracy Percentage of correct answers in the current run
reward Calculated reward score for the capability-evolver
trend Historical performance direction
focus_subjects List of specific topics requiring improvement
curriculum_stage The current level of the learning queue
next_questions Priority questions for the next iteration

HLE Benchmark Evolver Advanced Features

  • Multi-cycle pipeline support for continuous autonomous agent refinement.
  • Dynamic evaluation command injection for seamless integration with external shell scripts.
  • Automated interval management and timing between evolution cycles.
  • Granular curriculum signal generation for fine-tuned modality and subject focus.
  • Integrated template-based reporting for simulation when live data is unavailable.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*