DOCX Toolkit for Openclaw

A versatile toolkit for extracting structured text, tables, and deduplicated images from modern .docx and legacy .doc Word documents.

zacjiang
v1.0.0
Mar 5, 2026
0
1.3k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install docx-toolkit

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install docx-toolkit using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is DOCX Toolkit?

DOCX Toolkit is a powerful set of utilities designed to bridge the gap between static Word documents and data-driven workflows. By providing specialized scripts for text and table extraction, it allows developers to transform unstructured office files into clean, machine-readable formats. Whether you are dealing with modern XML-based .docx files or legacy OLE2 .doc formats, this addition to your Openclaw Skills library ensures high-fidelity data recovery, including full support for CJK (Chinese, Japanese, Korean) characters.

Beyond simple text recovery, the toolkit excels at media management. It features a sophisticated image extraction engine that automatically deduplicates files using MD5 hashing and filters out insignificant icons. For those integrating document data into vision-based AI workflows, the built-in compression tools significantly reduce token usage and processing costs. This makes it an essential component for any developer building complex document analysis pipelines within the Openclaw Skills ecosystem.

DOCX Toolkit Use Cases

  • AI-Powered Document Analysis: Extracting clean text and tables for LLM summarization or sentiment analysis using Openclaw Skills.
  • Legacy Data Migration: Converting old .doc archives into modern database or markdown formats.
  • Automated Image Auditing: Pulling all embedded media from corporate documents for compliance or content review.
  • Cost-Effective Vision Pipelines: Compressing extracted images before sending them to vision APIs to save on operational costs.
  • Multi-lingual Processing: Accurately extracting CJK text for global document processing tasks.

How DOCX Toolkit Works

  1. The user provides a path to a .docx or legacy .doc file within the environment.
  2. The toolkit selects the appropriate parser—python-docx for modern files or olefile for legacy WordDocument streams.
  3. Text is extracted while preserving structural integrity, specifically converting tables into pipe-delimited Markdown-style rows.
  4. If image extraction is triggered, the system scans for embedded media, applies MD5 deduplication, and filters by file size.
  5. Optional post-processing scripts can then batch-resize images to optimize them for API-based Openclaw Skills workflows.

DOCX Toolkit Setup

To integrate this skill into your workflow, ensure you have Python 3.6+ installed and set up the necessary dependencies:

pip3 install python-docx olefile Pillow

You can then run specific scripts for your needs:

# For text extraction
python3 scripts/extract_text.py input.docx output.txt

# For image extraction
python3 scripts/extract_images.py input.docx output_dir/

DOCX Toolkit Data Schema & Taxonomy

Component Format/Method Details
Text Content UTF-8 Plain Text Full paragraph extraction with CJK support
Tables Pipe-Delimited Structured rows (e.g.,
Images Original (PNG/JPG) Extracted with MD5-based deduplication
Metadata File System Sequential naming (img_001) and size-based filtering

DOCX Toolkit Advanced Features

  • Intelligent Image Deduplication: Uses MD5 hash comparison to ensure identical images are only saved once.
  • Legacy Format Support: Full extraction capabilities for the OLE2 .doc format which many modern libraries ignore.
  • Vision API Optimization: Includes batch resizing tools to reduce the footprint of images before AI processing in Openclaw Skills.
  • CJK Character Integrity: Specialized handling for Unicode text to ensure perfect extraction of Asian languages.
  • Configurable Filtering: Set minimum size thresholds (e.g., <5KB) to automatically skip UI icons and decorative elements.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*