ClawText Ingest for Openclaw

ClawText Ingest provides production-ready memory ingestion for AI agents, supporting Discord, files, and URLs with automatic deduplication.

ragesaq
v1.0.1
Mar 5, 2026
0
1.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install clawtext-ingest

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install clawtext-ingest using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is ClawText Ingest?

ClawText Ingest is a powerful developer tool designed to bridge the gap between unstructured data sources and AI agent memory. By leveraging Openclaw Skills, developers can transform Discord forums, local files, web URLs, and JSON data into structured, deduplicated memories. This skill ensures that AI agents have access to high-quality, context-aware information through automatic YAML frontmatter generation and project-based routing.

Built for reliability and scale, this tool solves common challenges such as manual data entry, duplicate memories, and the loss of hierarchy in complex platforms like Discord. As a core component of the Openclaw Skills ecosystem, it integrates seamlessly with RAG layers, allowing agents to fetch and process knowledge autonomously through documented interaction patterns.

ClawText Ingest Use Cases

  • Automating the ingestion of Discord forum posts and replies into an AI knowledge base.
  • Synchronizing local documentation or research files with an agent memory cluster.
  • Building autonomous RAG pipelines that periodically scrape URLs for updated information.
  • Managing team decisions and meeting notes by ingesting Slack exports or JSON files.
  • Integrating knowledge capture directly into agent workflows using direct API calls.

How ClawText Ingest Works

  1. The user identifies a data source such as a Discord forum, a set of Markdown files, or a specific URL.
  2. ClawText Ingest processes the source, applying SHA1-based cryptographic deduplication to ensure no redundant data is stored.
  3. The skill extracts content and generates rich metadata, preserving hierarchies like post-to-reply structures for Discord data.
  4. Memories are saved with YAML frontmatter, organizing them by project, type, and associated entities.
  5. The user or agent triggers a cluster rebuild, enabling the RAG layer to inject this new context into subsequent AI prompts.

ClawText Ingest Setup

To get started with this addition to your Openclaw Skills library, install the package via npm or the OpenClaw CLI:

npm install clawtext-ingest
# OR
openclaw install clawtext-ingest

For Discord ingestion, set your DISCORD_TOKEN environment variable and use the following command to fetch forum data:

# Inspect forum structure
clawtext-ingest-discord describe-forum --forum-id YOUR_FORUM_ID

# Ingest with progress
clawtext-ingest-discord fetch-discord --forum-id YOUR_FORUM_ID

Finally, sync your changes to the memory cluster to enable RAG functionality:

clawtext-ingest rebuild

ClawText Ingest Data Schema & Taxonomy

The skill organizes data into structured memory files with consistent metadata taxonomy. Below is the schema for generated memories:

Field Description
project The unique identifier for the knowledge category (e.g., 'docs', 'research').
type The content classification (e.g., fact, decision, thread, adr).
entities Auto-extracted keywords or linked concepts for agent retrieval.
date ISO-8601 timestamp of when the content was ingested.
hash SHA1 cryptographic hash used for 100% idempotent deduplication.

Cross-session tracking is maintained in a .ingest_hashes.json file to ensure that Openclaw Skills can track processed items across different runs.

ClawText Ingest Advanced Features

  • Support for 6 distinct agent patterns including Direct API, Discord Agent, and CLI Subprocesses.
  • Idempotent processing ensures that repeating the same command never creates duplicate memories.
  • Auto-batch mode for large Discord forums, switching to streaming for sets exceeding 500 posts.
  • Programmatic Node.js API allowing custom transforms and field mapping for complex integrations.
  • Real-time progress callbacks to monitor large-scale ingestion tasks within Openclaw Skills workflows.
  • Automatic cluster indexing to make new memories immediately available for RAG injection.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*