A versatile toolkit for extracting structured text, tables, and deduplicated images from modern .docx and legacy .doc Word documents.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install docx-toolkit
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install docx-toolkit using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
DOCX Toolkit is a powerful set of utilities designed to bridge the gap between static Word documents and data-driven workflows. By providing specialized scripts for text and table extraction, it allows developers to transform unstructured office files into clean, machine-readable formats. Whether you are dealing with modern XML-based .docx files or legacy OLE2 .doc formats, this addition to your Openclaw Skills library ensures high-fidelity data recovery, including full support for CJK (Chinese, Japanese, Korean) characters.
Beyond simple text recovery, the toolkit excels at media management. It features a sophisticated image extraction engine that automatically deduplicates files using MD5 hashing and filters out insignificant icons. For those integrating document data into vision-based AI workflows, the built-in compression tools significantly reduce token usage and processing costs. This makes it an essential component for any developer building complex document analysis pipelines within the Openclaw Skills ecosystem.
To integrate this skill into your workflow, ensure you have Python 3.6+ installed and set up the necessary dependencies:
pip3 install python-docx olefile Pillow
You can then run specific scripts for your needs:
# For text extraction
python3 scripts/extract_text.py input.docx output.txt
# For image extraction
python3 scripts/extract_images.py input.docx output_dir/
| Component | Format/Method | Details |
|---|---|---|
| Text Content | UTF-8 Plain Text | Full paragraph extraction with CJK support |
| Tables | Pipe-Delimited | Structured rows (e.g., |
| Images | Original (PNG/JPG) | Extracted with MD5-based deduplication |
| Metadata | File System | Sequential naming (img_001) and size-based filtering |
Loading
An AI-powered skill for automated bid and tender document auditing, cross-referencing requirements with responses to identify compliance risks.

A high-performance CLI for on-device Whisper transcription and Qwen3-TTS synthesis optimized for the Apple Neural Engine.

A lightweight monitoring tool that bridges AI agents with the ClawHQ dashboard for real-time status and task reporting.

A modular dashboard for orchestrating multiple OpenClaw agent brains with execution streams and automated task management.

A lightweight tool to generate professional PDF invoices from natural language or structured data without external services.

A lightweight Python-based tool to convert Markdown files into professional PDF documents with full CJK character support.








































