OpenOCR Skill for Openclaw

An intelligent OCR and document parsing engine that extracts text, mathematical formulas, and tables from images and PDFs with high precision.

topdu
v0.1.4
Feb 12, 2026
0
1.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install opencr-skill

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install opencr-skill using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is OpenOCR Skill?

OpenOCR is a comprehensive and efficient general OCR system designed to handle a wide variety of visual recognition tasks. As a key component in the Openclaw Skills ecosystem, it provides developers with a unified interface for text detection, recognition, and sophisticated document parsing. By leveraging Vision Language Models (VLM), it excels at recognizing complex structures like mathematical formulas and tables that traditional OCR systems often struggle with.

This skill is built for versatility, offering both lightweight mobile modes for speed and server-grade modes for maximum accuracy. Whether you are processing a single screenshot or a multi-page scanned PDF, OpenOCR provides the tools necessary to transform visual data into structured, machine-readable formats like Markdown or JSON, enhancing the capabilities of any Openclaw Skills implementation.

OpenOCR Skill Use Cases

  • Converting scanned PDF reports into structured Markdown for easier indexing and search.
  • Extracting LaTeX-formatted mathematical formulas from technical papers or screenshots.
  • Automating the digitization of printed tables into JSON data for analysis.
  • Performing layout analysis on complex documents to identify headings, paragraphs, and image captions.
  • Building automated data entry pipelines that process images from a directory using Openclaw Skills.

How OpenOCR Skill Works

  1. Input ingestion: The skill accepts image files (JPG, PNG, BMP) or PDF documents as the primary source.
  2. Task Selection: Users define the specific operation, such as detection (det), recognition (rec), end-to-end OCR (ocr), or universal recognition (unirec).
  3. Layout Analysis: For document tasks, the engine uses PP-DocLayoutV2 to identify different content blocks like text, tables, and charts.
  4. Recognition Pipeline: The chosen backend (ONNX or PyTorch) processes the visual data to extract text strings, formula LaTeX, or table structures.
  5. Result Serialization: The processed data is returned as a structured object, which can be saved as Markdown, JSON, or visualized with bounding boxes.

OpenOCR Skill Setup

To begin using this skill within your Openclaw Skills environment, install the package via pip. You can choose the basic installation or include optional dependencies for GPU support.

# Basic installation (CPU, ONNX backend)
pip install openocr-python

# GPU-accelerated ONNX inference
pip install openocr-python[onnx-gpu]

# PyTorch backend for high-accuracy server mode
pip install openocr-python[pytorch]

# Install all optional dependencies including Gradio demos
pip install openocr-python[all]

OpenOCR Skill Data Schema & Taxonomy

OpenOCR provides detailed metadata for every extraction task. The following table describes how the data is organized when utilizing Openclaw Skills:

Feature Data Type Description
Boxes np.ndarray Bounding box coordinates for detected text regions
Text String The recognized characters or LaTeX formulas
Score Float Confidence score (0.0 to 1.0) for the recognition
Elapse Float Time taken in seconds to process the task
Layout Blocks Dictionary Structured content identified during document parsing

OpenOCR Skill Advanced Features

  • Universal Recognition (UniRec): Uses a 0.1B parameter VLM to handle text, formulas, and tables in a single pass.
  • Multi-Backend Support: Switch seamlessly between ONNX for CPU efficiency and PyTorch for GPU-accelerated accuracy.
  • PDF Native Processing: Automatically handles multi-page PDF documents, converting pages to images internally for analysis.
  • Layout Analysis: Advanced document parsing that distinguishes between different types of content blocks for better structure retention.
  • Visualization Tools: Generate images with overlaid bounding boxes and labels to verify extraction accuracy.
  • CLI Integration: Execute complex OCR tasks directly from the command line, facilitating automation within Openclaw Skills.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*