Ghost Eye for Openclaw

Ghost Eye functions as an image preprocessing bridge, converting visual data into rich OCR text and scene descriptions for pure-text LLMs.

hunter-crk
v1.0.4
Jul 9, 2026
1
1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install ghost-eye

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install ghost-eye using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Ghost Eye?

Ghost Eye is a developer tool designed to bridge the gap between text-only language models and visual inputs. When an image is introduced to a conversation, this utility interceptively routes the file to an OpenAI-compatible vision model (such as Nex-N2-Pro) to execute OCR and generate descriptive summaries. The resulting plain-text metadata is then structured and appended to the LLM context, enabling text-based models to process image contents without requiring native multimodal APIs.

This approach helps teams avoid maintaining complex or expensive multimodal API pipelines for legacy models. Ghost Eye supports flexible configurations within developer workflows, integrating cleanly into multi-agent systems and custom Openclaw Skills environments.

Ghost Eye Use Cases

  • Extracting structural text and tables from uploaded document scans, screenshots, or receipts in real time.
  • Generating markdown-formatted descriptions of application interfaces, wireframes, or charts for code generation engines.
  • Lowering processing costs by offloading visual interpretation to specialized vision endpoints before passing descriptions to expensive reasoning models.
  • Facilitating visual comprehension for specialized local models that do not natively support visual input parameters.

How Ghost Eye Works

  1. Input Detection: The utility accepts image inputs provided via a local file path, a public URL, or a base64 encoded string.
  2. Preprocessing & Compression: The image undergoes format validation (supporting JPEG, PNG, WebP, GIF, or BMP) and is resized using Pillow to a maximum limit of 1920px on the longest edge at 85% quality to optimize transfer sizes.
  3. Caching Validation: Ghost Eye checks a local cache directory using an MD5 hash of the raw image bytes to verify if a prior analysis exists within the active TTL duration (default 7 days).
  4. Vision Processing API Call: On a cache miss, the system forwards the processed image to the configured OpenAI-compatible endpoint for structural and semantic analysis.
  5. Context Injection: The resulting OCR output and visual summary are structured in plain text and returned as a structured JSON object to be injected directly into the active LLM context.

Ghost Eye Setup

To add the capability to your setup, update your configuration in openclaw.json under skills.entries:

{
  "ghost-eye": {
    "enabled": true,
    "apiKey": { "source": "env", "provider": "default", "id": "NEXN2_API_KEY" },
    "env": {
      "NEXN2_BASE_URL": "https://api.siliconflow.cn/v1",
      "NEXN2_MODEL_NAME": "nex-agi/Nex-N2-Pro",
      "NEXN2_IMAGE_COMPRESS": "true",
      "NEXN2_CACHE_ENABLE": "true",
      "NEXN2_CACHE_TTL_DAYS": "7",
      "NEXN2_TIMEOUT_MS": "30000"
    }
  }
}

Ensure your NEXN2_API_KEY is exported to your runtime environment. To test the processing manually, call the analysis script directly:

python3 scripts/analyze.py --image-path "/absolute/path/to/image.png"

Ghost Eye Data Schema & Taxonomy

Environment Configuration Variables

Variable Type Required Default Description
NEXN2_API_KEY String Yes — Authentication key for your vision provider.
NEXN2_BASE_URL String No https://api.siliconflow.cn/v1 Target OpenAI-compatible API base URL.
NEXN2_MODEL_NAME String No nex-agi/Nex-N2-Pro Model designated for visual extraction.
NEXN2_CACHE_ENABLE Boolean No true Enables local result caching to minimize API usage.
NEXN2_CACHE_TTL_DAYS Integer No 7 Number of days until local cache files expire.

Executable Response Structure

Upon execution, the utility returns a structured JSON output representing the extraction status and content:

{
  "success": true,
  "content": "[Extracted OCR Text]\n\n[Scene and Visual Summary Description]",
  "metadata": {
    "model": "nex-agi/Nex-N2-Pro",
    "tokens_used": 1200,
    "cached": false,
    "process_time_ms": 1500
  }
}

Ghost Eye Advanced Features

  • Auto-Preprocessing Mode: Operates transparently within pipeline boundaries, automatically analyzing visual elements without requiring manual intervention from the end user.
  • Explicit Tool-Calling Mode: Allows the primary model to make selective tool queries (analyze_image_by_nexn2) when visual input processing is determined necessary.
  • Robust Image Optimization: Automatic validation of magic bytes and on-the-fly compression to protect the runtime against size-limit failures.
  • MD5 Hash Caching: Reduces cost and API latency on repeated file analysis by mapping files to local cached JSON payloads using a customizable TTL window.
  • Error Safety Boundaries: Network exceptions and formatting failures are encapsulated within JSON responses to prevent operational crashes in continuous workflows.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*