iWork to Markdown for Openclaw

iWork to Markdown is an Openclaw Skills parser that converts Apple Pages, Numbers, and Keynote files into clean Markdown text.

slearnai
v0.1.0
Jul 29, 2026
0
2.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install iwork2md

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install iwork2md using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is iWork to Markdown?

iWork to Markdown is an Openclaw Skills utility for extracting readable content from Apple iWork documents, including .pages, .numbers, and .key files. It is designed for when you need to read, search, summarize, or repurpose the contents of a document without opening Apple’s apps.

This skill focuses on reliable text recovery and Markdown generation rather than pixel-perfect rendering. It uses a dependency-free Python implementation to unpack iWork bundles, decode the underlying .iwa data, recover text fragments, reconstruct simple table structure from Numbers, and expose embedded media references in a developer-friendly format.

iWork to Markdown Use Cases

  • Convert a Pages document into Markdown for editing, analysis, or publishing.
  • Extract text from a Numbers spreadsheet without manually copying cells.
  • Read Keynote slide content as plain Markdown for review or summarization.
  • Inspect embedded media assets in an iWork file.
  • Rapidly translate iWork content into text for downstream AI workflows.
  • Use Openclaw Skills when a user asks to "open", "convert", or "extract" content from an iWork document.

How iWork to Markdown Works

  1. Opens the iWork package and locates Index.zip or the raw .iwa files inside the document bundle.
  2. Removes iWork’s non-standard Snappy framing from each .iwa record and decompresses the payload using raw Snappy logic.
  3. Parses the Protobuf archive container generically, walking message fields to recover UTF-8 string content without requiring an app-specific schema map.
  4. Reconstructs Markdown output by deriving a title, listing embedded media, grouping Numbers row strings into tables, and collecting the remaining text fragments into readable sections.
  5. Writes the result to a .md file, prints to stdout, or exposes debug/media inspection modes for deeper analysis.

iWork to Markdown Setup

The skill is dependency-free and runs on Python 3.8+ using only the standard library.

  1. Ensure Python 3.8 or newer is available.
  2. Run the converter against an iWork file.
python3 scripts/iwork2md.py path/to/Doc.pages
  1. Optionally choose an explicit output path.
python3 scripts/iwork2md.py Doc.numbers out.md
  1. Print Markdown to stdout for piping into other tools.
python3 scripts/iwork2md.py Doc.key --stdout
  1. Use debug flags when you need inspection or validation.
python3 scripts/iwork2md.py Doc.numbers --texts
python3 scripts/iwork2md.py Doc.pages --media
python3 scripts/test_iwa.py

iWork to Markdown Data Schema & Taxonomy

iWork to Markdown organizes output around the recovered document content rather than the original visual layout.

Component Purpose Notes
Source bundle Input .pages, .numbers, or .key package Contains Index.zip or direct .iwa files
.iwa records Internal iWork data blocks Stored with non-standard Snappy framing
ArchiveInfo / Protobuf payloads Generic message container Walked to recover string fields
Metadata/Properties.plist Document metadata source Used for title when available
Markdown body Primary output artifact Written to .md or stdout
Embedded media list Inventory of images/video Reported separately from prose
Numbers table reconstruction Tabular text recovery Consecutive `cell

Metadata taxonomy used by the converter:

  • Title: from Metadata/Properties.plist or inferred from the first heading.
  • Body text: largest multi-line recovered text block.
  • Fragments: remaining UTF-8 text strings not promoted into the body.
  • Media: embedded image/video references found in the bundle.
  • Tables: row strings normalized into Markdown table syntax when detected.

Limitations reflected in the schema:

  • No exact layout, typography, chart, or shape fidelity.
  • Encrypted or password-protected documents are not supported.
  • Structural recovery is content-first, not WYSIWYG.

iWork to Markdown Advanced Features

  • Dependency-free implementation with no third-party Python packages.
  • Supports Apple Pages, Numbers, and Keynote documents in one workflow.
  • Generic Protobuf walking recovers readable text without a version-specific schema map.
  • Reconstructs Numbers tables into Markdown from serialized row strings.
  • Debug mode can dump every recovered text fragment for forensic inspection.
  • Media inventory mode lists embedded images and video assets.
  • Stdout output makes it easy to chain Openclaw Skills into shell pipelines and AI automation.
  • Includes a synthetic round-trip test script to validate Snappy framing and parser behavior.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*