A robust data extraction tool designed to turn web pages into structured JSON, CSV, or Markdown datasets using CSS selectors and auto-detection.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install smart-web-scraper
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install smart-web-scraper using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The Smart Web Scraper is a developer-centric tool built to simplify the process of gathering data from the web. By utilizing high-performance parsing libraries like BeautifulSoup4 and lxml, it allows users to pinpoint specific data points using CSS selectors or automatically identify complex structures like pricing tables and paginated lists. This skill bridges the gap between raw HTML and actionable data, providing a clean interface for data scientists and developers alike.
As a core component of the Openclaw Skills library, this scraper is optimized for speed and reliability. It handles the nuances of web requests, including redirects and encoding detection, while ensuring that the output is perfectly formatted for downstream AI analysis or database ingestion.
To use this skill within the Openclaw Skills environment, ensure you have the uv package manager installed. You can run commands directly using the following pattern:
# Basic extraction
uv run --with beautifulsoup4 --with lxml python scripts/scraper.py extract "https://example.com"
# Extraction with CSS selectors and JSON output
uv run --with beautifulsoup4 --with lxml python scripts/scraper.py extract "https://example.com/products" -s ".item" -f json
The skill organizes extracted data into a predictable format. Below is the schema for standard extraction results:
| Property | Type | Description |
|---|---|---|
text |
string | The inner text of the captured element |
tag |
string | The HTML tag name (e.g., div, p, span) |
class |
string | The CSS classes associated with the element |
href |
string | The URL (only present when using the links command) |
metadata |
object | Contains page title, meta descriptions, and headers when using the structure command |
Loading
A natural language cron scheduler that allows users to automate tasks using simple English commands instead of complex cron syntax.

A lightweight service and endpoint monitoring utility for tracking the health of self-hosted infrastructure and SSL certificates.

A unified CLI dashboard to track daily earnings, payouts, and node uptime across multiple bandwidth-sharing and crypto-passive income applications.

A unified monitoring tool for tracking multiple passive income streams, including bandwidth nodes, decentralized storage, and DeFi staking.

A proactive security monitoring agent that identifies vulnerabilities in Docker images and software dependencies while providing real-time CVE alerting.

A comprehensive monitoring tool for Web3 grant rounds, DAO funding, and quadratic funding matching estimates.








































