Smart Web Scraper for Openclaw

A robust data extraction tool designed to turn web pages into structured JSON, CSV, or Markdown datasets using CSS selectors and auto-detection.

mariusfit
v1.0.0
Feb 24, 2026
0
3.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install smart-web-scraper

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install smart-web-scraper using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Smart Web Scraper?

The Smart Web Scraper is a developer-centric tool built to simplify the process of gathering data from the web. By utilizing high-performance parsing libraries like BeautifulSoup4 and lxml, it allows users to pinpoint specific data points using CSS selectors or automatically identify complex structures like pricing tables and paginated lists. This skill bridges the gap between raw HTML and actionable data, providing a clean interface for data scientists and developers alike.

As a core component of the Openclaw Skills library, this scraper is optimized for speed and reliability. It handles the nuances of web requests, including redirects and encoding detection, while ensuring that the output is perfectly formatted for downstream AI analysis or database ingestion.

Smart Web Scraper Use Cases

  • Gathering product pricing and availability from e-commerce sites.
  • Extracting financial tables and research data for analytical reporting.
  • Crawling blogs and news sites to archive articles or follow pagination.
  • Mapping website architecture by extracting internal and external link networks.

How Smart Web Scraper Works

  1. The user initiates a request specifying a URL and a desired action such as extract, tables, or links.
  2. The system fetches the page content using a standard browser User-Agent to ensure compatibility.
  3. The parser processes the HTML DOM based on provided CSS selectors or auto-detection logic for tables.
  4. Data is structured into the requested format (JSON, CSV, or Markdown).
  5. The tool outputs the result to the console or saves it directly to a file while respecting rate limits.

Smart Web Scraper Setup

To use this skill within the Openclaw Skills environment, ensure you have the uv package manager installed. You can run commands directly using the following pattern:

# Basic extraction
uv run --with beautifulsoup4 --with lxml python scripts/scraper.py extract "https://example.com"

# Extraction with CSS selectors and JSON output
uv run --with beautifulsoup4 --with lxml python scripts/scraper.py extract "https://example.com/products" -s ".item" -f json

Smart Web Scraper Data Schema & Taxonomy

The skill organizes extracted data into a predictable format. Below is the schema for standard extraction results:

Property Type Description
text string The inner text of the captured element
tag string The HTML tag name (e.g., div, p, span)
class string The CSS classes associated with the element
href string The URL (only present when using the links command)
metadata object Contains page title, meta descriptions, and headers when using the structure command

Smart Web Scraper Advanced Features

  • Intelligent multi-page crawling that follows pagination links to a specified depth.
  • Automated HTML table detection that converts complex grids into clean CSV rows.
  • Respectful crawling mechanisms including configurable delays and robots.txt validation.
  • Support for multiple output formats including Markdown for seamless documentation integration.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*