DocStrange (Nanonets Document Extraction) for Openclaw

DocStrange is a powerful document extraction API by Nanonets that converts PDFs and images into structured Markdown, JSON, or CSV data with field-level confidence scoring.

shhdwi
v1.0.2
Feb 11, 2026
21
5.3k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install docstrange

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install docstrange using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is DocStrange (Nanonets Document Extraction)?

DocStrange by Nanonets is a specialized document intelligence tool designed to bridge the gap between unstructured physical documents and structured digital data. By utilizing advanced OCR and machine learning models, it allows developers to parse invoices, receipts, and complex forms with high accuracy. As part of the Openclaw Skills ecosystem, DocStrange enables AI agents to process document layouts, extract tabular data, and verify information through confidence scores, making it an essential component for automated data entry and document analysis pipelines.

This skill supports both synchronous processing for quick tasks and asynchronous workflows for large, multi-page documents. Whether you are building a financial automation bot or a research assistant that needs to ingest whitepapers, integrating this into your Openclaw Skills setup provides a robust way to handle file-based data extraction without manual intervention.

DocStrange (Nanonets Document Extraction) Use Cases

  • Automating the extraction of line items and totals from invoices and receipts for accounting software.
  • Converting scanned PDF reports into clean Markdown for indexing in knowledge bases.
  • Parsing bank statements and financial tables directly into CSV format for data analysis.
  • Digitizing historical archives or legal contracts into structured JSON using custom schemas.
  • Enhancing RAG (Retrieval-Augmented Generation) pipelines by providing high-quality text extraction from complex layouts.

How DocStrange (Nanonets Document Extraction) Works

  1. The skill receives a document input via a file upload, public URL, or base64 encoded string.
  2. It sends a request to the Nanonets extraction API with specific instructions on the desired output format (Markdown, JSON, or CSV).
  3. For structured JSON data, the skill can apply a predefined schema to ensure the extracted fields match your application's data requirements.
  4. The document is processed using OCR and layout engines to identify text, tables, and key-value pairs.
  5. The skill returns the structured content along with metadata, such as bounding boxes or confidence scores for every extracted field.
  6. For documents exceeding five pages, the skill manages an asynchronous lifecycle by queuing the file and polling for the final result.

DocStrange (Nanonets Document Extraction) Setup

To use this skill, first obtain an API key from the Nanonets dashboard at https://docstrange.nanonets.com/app. Then, configure your environment and Openclaw Skills settings:

# Set your API key as an environment variable
export DOCSTRANGE_API_KEY="your_api_key_here"

Add the skill to your configuration file:

{
  "skills": {
    "entries": {
      "docstrange": {
        "enabled": true
      }
    }
  }
}

For security, ensure your configuration file permissions are restricted using chmod 600 ~/.openclaw/openclaw.json.

DocStrange (Nanonets Document Extraction) Data Schema & Taxonomy

DocStrange organizes extracted data into structured formats based on the requested output type. It supports the following schema structures:

Feature Format Description
Text Content Markdown Preserves headers, bold text, and basic lists for readability.
Structured Data JSON Returns key-value pairs. Can be guided by a JSON Schema for strict typing.
Tabular Data CSV Specifically targets tables within the document for spreadsheet compatibility.
Metadata Object Includes confidence_scores (0-100) and bounding_boxes for layout analysis.

Example JSON schema for an invoice would include properties for invoice_number, date, vendor, and an array of line_items.

DocStrange (Nanonets Document Extraction) Advanced Features

  • Bounding Box Metadata: Retrieve exact coordinates for every extracted element to perform detailed layout analysis.
  • Hierarchy Output: Extract complex document structures including nested sections and parent-child relationships.
  • Financial Documents Mode: Specialized processing for accounting documents to improve currency and table formatting.
  • Custom Extraction Instructions: Provide natural language prompts to guide the AI on what specific data to prioritize or ignore.
  • Confidence Scoring: Receive field-level reliability metrics to trigger human-in-the-loop reviews for low-confidence data.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*