Agent Browser for Openclaw

A specialized browser automation CLI that allows AI agents to interact with web pages through a unique element-referencing system.

chulla-ceja
v0.1.0
Mar 16, 2026
0
1.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install agent-browser-6

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install agent-browser-6 using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Agent Browser?

Agent Browser is a high-performance CLI utility built for AI agents to programmatically navigate the web. By utilizing Chrome/Chromium via the Chrome DevTools Protocol (CDP), it provides a structured interface for agents to perform tasks like clicking buttons, filling forms, and extracting data. This tool is essential for developers building sophisticated Openclaw Skills that require reliable, headless, or headed web interaction.

Unlike traditional automation frameworks, Agent Browser is optimized for agentic workflows, featuring a snapshot system that maps the page's accessibility tree into simple element references. This allows agents to reason about the UI and interact with specific elements without complex CSS selectors, making it a cornerstone for modern Openclaw Skills development.

Agent Browser Use Cases

  • Automating complex web form submissions and multi-step registration flows.
  • Scrapping data from dynamic, JavaScript-heavy websites that require user interaction.
  • Performing visual regression testing and snapshot diffing for web applications.
  • Creating automated walkthroughs and documentation using annotated screenshots.
  • Managing authenticated sessions across different domains for secure data retrieval.

How Agent Browser Works

  1. Initialization: The agent opens a session and navigates to a specific URL using the open command.
  2. Element Mapping: A snapshot is captured to generate an interactive map of the page, assigning unique references like @e1 to elements.
  3. Programmatic Interaction: The agent executes commands such as click, fill, or select using the generated references.
  4. Dynamic Re-Snapshotting: After navigation or DOM updates, the agent takes a new snapshot to refresh element references and maintain state accuracy.
  5. Outcome Verification: The agent uses diffing tools or screenshots to verify that the performed actions achieved the desired results.

Agent Browser Setup

Install Agent Browser via your preferred package manager to start building your Openclaw Skills:

npm i -g agent-browser
# Or using Homebrew
brew install agent-browser
# Or using Cargo
cargo install agent-browser

# Initialize the browser engine
agent-browser install

To enable encryption for session data, set the AGENT_BROWSER_ENCRYPTION_KEY environment variable in your terminal configuration.

Agent Browser Data Schema & Taxonomy

Agent Browser organizes its operational data using structured JSON files and standard media formats to ensure compatibility with various Openclaw Skills.

Data Component Format Purpose
Session State .json Stores cookies, localStorage, and session tokens.
Page Snapshot Text/JSON Maps the accessibility tree to interactive references (@e1, @e2).
Screenshots .png / .jpg Visual captures of the page for debugging or vision-based agents.
Auth Vault Encrypted Binary Securely stores site credentials using AES encryption.
Action Policy .json Defines allowed and denied commands for security gating.

Agent Browser Advanced Features

  • Annotated Vision Mode: Generates screenshots with numbered labels overlaid on interactive elements for spatial reasoning.
  • Lightpanda Engine Support: Switch to a lightweight headless engine for 10x faster execution in performance-critical Openclaw Skills.
  • Security Content Boundaries: Wraps untrusted page content in cryptographic nonces to prevent prompt injection attacks.
  • iOS Simulator Integration: Automate mobile Safari on macOS using Appium and Xcode for mobile-specific web tasks.
  • Command Chaining: Execute multiple interactions in a single shell invocation using standard operators for maximum efficiency.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*