Table OCR for Openclaw

A high-performance OCR tool that extracts complex table structures from scanned documents and images into structured Markdown.

mzlzyca
v0.4.0
Apr 3, 2026
0
1.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install table-ocr

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install table-ocr using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Table OCR?

Table OCR is a specialized tool powered by MinerU from OpenDataLab, designed to bridge the gap between static image-based documents and actionable data. It leverages advanced table detection and recognition to handle complex layouts, including merged cells and multi-lingual content. By integrating this into your collection of Openclaw Skills, you can automate the digitization of printed reports, financial statements, and screenshots.

This skill is particularly effective for developers and data scientists who need to transform unstructured PDF or image data into structured Markdown or CSV-ready formats. Whether you are dealing with a grainy photo of a spreadsheet or a multi-page scanned report, Table OCR ensures high fidelity in structural preservation and text accuracy.

Table OCR Use Cases

  • Automating data entry from scanned financial reports and invoices.
  • Converting table screenshots into structured Markdown for documentation.
  • Digitizing legacy printed documents into machine-readable formats for LLM processing.
  • Extracting data from multi-page PDFs with complex, nested table layouts.
  • Processing image-based tables in English, Chinese, and other supported languages.

How Table OCR Works

  1. The user provides a local file path or a URL for a PDF or image file (.png, .jpg, .jpeg, .webp).
  2. The skill authenticates with the MinerU API using the provided MINERU_TOKEN.
  3. The system performs an initial pass to detect table boundaries and structural elements like merged cells.
  4. Optical Character Recognition (OCR) is applied to extract text from the identified table cells.
  5. The tool reconstructs the table structure and outputs the result as structured Markdown or saves it to a specified directory.

Table OCR Setup

To use this within your Openclaw Skills workflow, install the MinerU CLI and configure your authentication token.

# Install via npm
npm install -g mineru-open-api

# Or install via Go (macOS/Linux)
go install github.com/opendatalab/MinerU-Ecosystem/cli/mineru-open-api@latest

# Set up your authentication token
mineru-open-api auth
# Or export directly
export MINERU_TOKEN="your-token-here"

Table OCR Data Schema & Taxonomy

The skill manages inputs and outputs through the following schema:

Feature Specification
Supported Inputs .pdf, .png, .jpg, .jpeg, .webp
Output Formats Markdown (stdout), DOCX (requires -o flag)
Default Language Chinese (ch), English (en) via --language flag
Data Organization Metadata is stored via openclaw tags; output files are saved to user-defined directories using the -o flag.
API Dependency Requires valid token from mineru.net

Table OCR Advanced Features

  • Multi-lingual support with explicit language hints using the --language flag.
  • Targeted extraction by specifying page ranges with the --pages parameter for large PDF documents.
  • Combined OCR and table detection in a single pass for maximum efficiency.
  • Support for both local file system processing and direct URL ingestion.
  • Seamless integration with AI agents as part of the broader Openclaw Skills ecosystem for automated workflows.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Requires
Bins mineru-open-api
Github Stars: 0
forks: 0

Featured*