Substack Scraper for Openclaw

A high-performance scraping tool designed to extract articles, metadata, and engagement metrics from Substack publications via the Apify API.

marcindudekdev
v1.0.0
Mar 7, 2026
0
866
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install apify-substack-scraper

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install apify-substack-scraper using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Substack Scraper?

The Substack Scraper is a robust integration designed for developers and researchers who need to programmatically access newsletter content. By utilizing an Apify Actor, this skill allows users to crawl Substack publications to retrieve full article text, author information, and publication dates. As part of the wider ecosystem of Openclaw Skills, it provides a seamless way to feed niche newsletter data into AI agents or content pipelines.

This tool is particularly valuable for those looking to archive digital content or perform competitive analysis within the newsletter space. It handles the complexities of web crawling and data structuring, delivering clean JSON output that is ready for analysis or redistribution. By incorporating Openclaw Skills into your workflow, you can bypass manual data collection and focus on deriving insights from the scraped content.

Substack Scraper Use Cases

  • Monitoring specific Substack publications for new industry trends and updates.
  • Archiving long-form newsletter content for offline reading or personal knowledge bases.
  • Collecting engagement data and article metadata for market research and sentiment analysis.
  • Automating the ingestion of newsletter feeds into specialized Openclaw Skills for summarization.

How Substack Scraper Works

  1. The user defines the target Substack publication URLs and sets parameters such as the maximum number of articles to retrieve.
  2. The skill validates the presence of the required APIFY_TOKEN and system utilities like curl and jq.
  3. A request is dispatched to the Apify REST API to trigger the scraper actor (ID: BULaGFURBV7WG3K81).
  4. The skill manages the execution, supporting both synchronous runs for small tasks and asynchronous polling for larger datasets.
  5. Upon completion, the dataset is fetched and parsed into a structured format for the user to review or export.

Substack Scraper Setup

To get started with this skill, ensure you have your Apify API credentials ready. Follow these steps to configure your environment:

  1. Obtain your API token from the Apify Console.
  2. Set the environment variable in your terminal:
export APIFY_TOKEN=your_token_here
  1. Verify that curl and jq are installed on your local machine to handle API requests and JSON parsing.
  2. When using this within the context of Openclaw Skills, ensure the primary environment variable is mapped correctly in your agent configuration.

Substack Scraper Data Schema & Taxonomy

The skill processes and returns data based on the following schema:

Parameter Type Description
urls Array A list of Substack publication URLs to be scraped.
maxArticles Integer The maximum number of posts to extract per publication.
includeContent Boolean If true, the full body text of the articles will be included.

Output Format: The result is a JSON array of items containing fields such as title, author, canonical_url, post_date, and body_html (if content is requested).

Substack Scraper Advanced Features

  • Support for asynchronous execution to handle large-scale scraping jobs without timeouts.
  • Granular control over the depth of the scrape via maxArticles constraints.
  • Ability to toggle full-text extraction to save on processing time and data volume.
  • Integration ready for downstream Openclaw Skills to perform automated summarization or translation.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Requires
Bins curljq
Github Stars: 0
forks: 0

Featured*