PDF Reader Skill for Openclaw

A high-performance text extraction tool designed to bridge the gap between static PDF documents and AI-ready data structures.

nantes
v1.0.3
Feb 22, 2026
2
1.4k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install pdfreader

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install pdfreader using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is PDF Reader Skill?

The PDF Reader skill is a specialized utility that leverages the PyMuPDF library to parse and extract text from PDF files. It is designed to empower AI agents with the ability to ingest complex documents, making it an essential component within the ecosystem of Openclaw Skills. By converting binary document data into structured text or JSON, it facilitates deeper analysis and searchability for any automated workflow.

This skill prioritizes performance and security, ensuring that text extraction is handled locally and within strict file system boundaries. Whether you are dealing with academic papers, financial reports, or technical manuals, this tool provides the raw data needed for intelligent processing.

PDF Reader Skill Use Cases

  • Automating the ingestion of research papers for AI-powered summarization
  • Converting legacy PDF reports into JSON format for data analysis pipelines
  • Extracting document metadata to organize large digital libraries
  • Feeding specific page ranges into LLMs to reduce token consumption and improve context accuracy

How PDF Reader Skill Works

  1. The user provides a target PDF file path and an optional page limit via the command line.
  2. The skill validates the file extension and ensures the path is within the permitted working directory.
  3. Using the PyMuPDF engine, the script opens the document and extracts the specified range of pages.
  4. Document metadata, including title, author, and creation date, is collected to provide context.
  5. The extracted text is cleaned of encoding artifacts and either displayed in the terminal or saved to a structured JSON file.

PDF Reader Skill Setup

To utilize this skill, you must first install the PyMuPDF dependency. Run the following command in your terminal:

pip install pymupdf

Ensure that the pdf_reader.py file is located in your active project directory. You can then execute the skill using standard Python commands, such as python pdf_reader.py "document.pdf" 10.

PDF Reader Skill Data Schema & Taxonomy

The skill produces a structured JSON output when the --output flag is used. This allows other Openclaw Skills to easily consume the data:

Key Type Description
source string The name of the original PDF file
metadata object Contains title, author, and document properties
pages array An array of strings containing the text for each extracted page
total_pages integer The number of pages successfully processed

PDF Reader Skill Advanced Features

  • Partial extraction support allows users to target specific page counts to optimize processing speed
  • Built-in security layer prevents path traversal attacks by restricting operations to the current directory
  • Intelligent encoding handling ensures that special characters and symbols are correctly preserved for AI reading
  • Seamless JSON serialization enables direct integration with multi-agent orchestration frameworks

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*