DocStrange is a powerful document extraction API by Nanonets that converts PDFs and images into structured Markdown, JSON, or CSV data with field-level confidence scoring.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install docstrange
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install docstrange using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
DocStrange by Nanonets is a specialized document intelligence tool designed to bridge the gap between unstructured physical documents and structured digital data. By utilizing advanced OCR and machine learning models, it allows developers to parse invoices, receipts, and complex forms with high accuracy. As part of the Openclaw Skills ecosystem, DocStrange enables AI agents to process document layouts, extract tabular data, and verify information through confidence scores, making it an essential component for automated data entry and document analysis pipelines.
This skill supports both synchronous processing for quick tasks and asynchronous workflows for large, multi-page documents. Whether you are building a financial automation bot or a research assistant that needs to ingest whitepapers, integrating this into your Openclaw Skills setup provides a robust way to handle file-based data extraction without manual intervention.
To use this skill, first obtain an API key from the Nanonets dashboard at https://docstrange.nanonets.com/app. Then, configure your environment and Openclaw Skills settings:
# Set your API key as an environment variable
export DOCSTRANGE_API_KEY="your_api_key_here"
Add the skill to your configuration file:
{
"skills": {
"entries": {
"docstrange": {
"enabled": true
}
}
}
}
For security, ensure your configuration file permissions are restricted using chmod 600 ~/.openclaw/openclaw.json.
DocStrange organizes extracted data into structured formats based on the requested output type. It supports the following schema structures:
| Feature | Format | Description |
|---|---|---|
| Text Content | Markdown | Preserves headers, bold text, and basic lists for readability. |
| Structured Data | JSON | Returns key-value pairs. Can be guided by a JSON Schema for strict typing. |
| Tabular Data | CSV | Specifically targets tables within the document for spreadsheet compatibility. |
| Metadata | Object | Includes confidence_scores (0-100) and bounding_boxes for layout analysis. |
Example JSON schema for an invoice would include properties for invoice_number, date, vendor, and an array of line_items.
Loading
A fully local, CPU-based text-to-speech engine providing high-quality audio generation and voice cloning without internet dependencies.

A multi-step deployment agent that orchestrates the workflow from local build and testing to GitHub hosting and Cloudflare Pages deployment.

A reputation-only protocol for AI agents providing cryptographic identity, peer attestation, and fast verification.

A robust multi-channel file delivery system designed to bypass sandbox and email restrictions through automated fallback logic and integrity verification.

A powerful, API-free internet search tool for AI agents to retrieve real-time data from Baidu and Bing.

A specialized tool for understanding and summarizing local PDF, video, and audio files using the Google Gemini API.








































