A local, privacy-focused tool that converts image-based PDFs into searchable text using Tesseract OCR and Python.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install pdf-ocr-extraction
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install pdf-ocr-extraction using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The PDF OCR Extractor is a specialized utility designed for the Openclaw Skills ecosystem to bridge the gap between static image scans and machine-readable text. Unlike many cloud-based alternatives, this skill operates entirely on your local machine, ensuring that sensitive documents are never transmitted to third-party servers.
By leveraging the Tesseract engine and Python's imaging libraries, it renders PDF pages into high-resolution images and applies advanced optical character recognition. This makes it an essential tool for developers and researchers who need to process historical archives, scanned receipts, or non-searchable document batches within their AI workflows.
To utilize this within the Openclaw Skills framework, ensure your system has the necessary binaries and Python packages installed:
# Install Tesseract OCR (Debian/Ubuntu example)
sudo apt-get install tesseract-ocr tesseract-ocr-eng tesseract-ocr-chi-sim
# Install required Python dependencies
pip install pypdfium2 pytesseract Pillow
The skill manages data through a temporary processing lifecycle to minimize storage footprint:
| Data Type | Description | Persistence |
|---|---|---|
| Input PDF | The source file provided by the user. | User-managed |
| Temp Images | Intermediate PNG files created per page in /tmp/. | Deleted after OCR |
| Extracted Text | Raw string output containing the OCR results. | Returned to Agent |
| Language Packs | Tesseract .traineddata files required for specific scripts. | System-installed |
Loading
A local, zero-dependency reference tool providing instant access to Page Saver documentation, patterns, and best practices directly in your terminal.

A dedicated New Zealand tax assistant that transforms receipt photos into IRD-ready GST reports and comprehensive annual tax summaries.

A powerful utility to send formatted HTML emails and notifications through the Resend API via a simple shell interface.

A zero-dependency reference tool providing instant access to Amortize devtools documentation and implementation patterns.

A professional-grade research agent that enforces data integrity through mandatory source carding and evidence-based reporting.

A specialized tool for extracting and organizing public metadata, membership details, and creator guidelines from the Disney Plus platform.








































