Smart OCR for Openclaw

A high-performance OCR skill for extracting text from images, screenshots, and scanned documents in over 100 languages using PaddleOCR.

duykhangdangzn1
v1.0.0
Feb 22, 2026
0
2.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install smar

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install smar using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Smart OCR?

Smart OCR is a robust text-extraction tool designed for the Openclaw Skills framework. It leverages the state-of-the-art PaddleOCR engine to identify and convert visual text into machine-readable data. Whether you are dealing with complex business cards, handwritten notes, or multi-page scanned PDFs, this skill provides the precision needed for modern AI workflows.

By integrating this tool into your Openclaw Skills setup, you gain access to features like angle classification, layout reconstruction, and multilingual support. It is optimized for both CPU and GPU environments, ensuring fast processing times for single images or large batches of documents.

Smart OCR Use Cases

  • Digitizing text from screenshots or business card photos.
  • Extracting content from multi-page scanned PDF documents.
  • Automating receipt scanning and financial data entry.
  • Analyzing text from multilingual documents containing languages like Chinese, Japanese, or Arabic.
  • Preprocessing visual data for search indexing within Openclaw Skills.

How Smart OCR Works

  1. The user provides an image file, PDF, or URL to the Openclaw Skills agent.
  2. The skill optionally detects the document language or uses a pre-specified language code.
  3. PaddleOCR processes the input to identify text blocks, bounding boxes, and confidence scores.
  4. Advanced layout reconstruction logic organizes the raw text into logical lines and paragraphs.
  5. The agent returns the structured text data or a summary to the user.

Smart OCR Setup

To get started with this skill in your Openclaw Skills environment, install the necessary dependencies:

# Install the OCR engine (CPU version)
pip install paddlepaddle paddleocr

# For GPU acceleration (optional)
pip install paddlepaddle-gpu

# Install additional document processing tools
pip install pdf2image Pillow

Smart OCR Data Schema & Taxonomy

The Smart OCR skill organizes its output into structured data, tracking spatial coordinates and text accuracy.

Attribute Description
text The actual string extracted from the image.
confidence A decimal score (0.0 - 1.0) representing extraction accuracy.
bbox Bounding box coordinates (left, top, right, bottom).
raw_box The four-point polygon vertices of the text area.
language The detected or specified language code used for OCR.

Smart OCR Advanced Features

  • Angle Classification: Automatically detects and corrects rotated text for higher accuracy.
  • Layout Reconstruction: Groups disparate text blocks into logical, readable paragraphs based on vertical and horizontal proximity.
  • Image Preprocessing: Built-in support for contrast enhancement and sharpening using Pillow to handle low-quality scans.
  • Batch Processing: Parallel execution support for high-throughput document analysis within Openclaw Skills.
  • Multilingual Fusion: Ability to merge results from different language models for complex documents.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*