Automates HLE benchmark reward ingestion and curriculum generation for AI capability evolution.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install hle-benchmark-evolver
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install hle-benchmark-evolver using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The HLE Benchmark Evolver is a specialized utility designed to operationalize Humanity's Last Exam (HLE) score-driven evolution. As part of the ecosystem of Openclaw Skills, it bridges the gap between raw benchmark results and actionable agent improvements. It allows developers to ingest question-level output and transform it into high-fidelity reward signals and curriculum queues, ensuring the agent focuses on the most impactful learning paths first.
This skill is particularly effective for teams looking to systematically increase their benchmark scores through automated feedback loops. By integrating with the capability-evolver, it provides a structured approach to model refinement, prioritizing easy-first curriculum stages and identifying specific focus subjects or modalities that require attention.
To get started with this component of Openclaw Skills, ensure you have your HLE benchmark reports ready. Run the basic ingestion using the following command:
node skills/hle-benchmark-evolver/run_result.js --report=/absolute/path/hle_report.json
For a full automated evolution cycle, use the pipeline script:
node skills/hle-benchmark-evolver/run_pipeline.js --report=/absolute/path/hle_report.json --cycles=1
You can also let the pipeline generate reports by providing an evaluation command:
node skills/hle-benchmark-evolver/run_pipeline.js --report=/path/hle_report.json --eval_cmd="python /path/to/eval_hle.py --out {{report}}" --cycles=3
The skill outputs a structured JSON contract containing the following metrics to track evolution progress:
| Field | Description |
|---|---|
| benchmark_id | Identifier for the benchmark (e.g., cais/hle) |
| accuracy | Percentage of correct answers in the current run |
| reward | Calculated reward score for the capability-evolver |
| trend | Historical performance direction |
| focus_subjects | List of specific topics requiring improvement |
| curriculum_stage | The current level of the learning queue |
| next_questions | Priority questions for the next iteration |
Loading
A streamlined utility for visualizing directory hierarchies in both human-readable ASCII and machine-readable JSON formats.

A meta-analysis tool that examines AI evolution memory graphs to detect stagnation and optimize success rates.

A lightweight sentiment analysis tool that automatically suggests contextually appropriate emoji reactions to increase engagement in digital conversations.

CritPt Solver provides a robust environment for validating and executing Python-based solutions for CritPt benchmark challenges.

A specialized tool for wrapping Humanity's Last Exam (HLE) questions in a structured Chain-of-Thought reasoning framework.

A lightweight utility for recursively scanning workspaces to identify and report syntax errors in JSON files.








































