Vision for Openclaw

Equips non-vision models like DeepSeek with the ability to see, analyze, and describe local or remote images using external vision APIs.

guorui999
v0.1.0
Jun 9, 2026
1
955
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install vision-2

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install vision-2 using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Vision?

The Vision skill is designed to bring sight to AI models that do not natively support vision capabilities, such as DeepSeek. By leveraging external, OpenAI-compatible vision APIs, this tool enables text-only models to process images, read text from screenshots, and perform complex visual reasoning. This capability is exceptionally valuable when integrated into larger workflows leveraging Openclaw Skills, allowing agents to understand UI layouts, parse diagrams, and compare mockups.

Whether you are asking the agent to analyze an error screenshot, extract text from an uploaded receipt, or compare different visual assets, this skill bridges the gap seamlessly. It converts local files or online image URLs into structured textual descriptions that your primary language model can easily interpret.

Vision Use Cases

  • Visual Troubleshooting: Send a screenshot of an error message or terminal output, and let the agent diagnose the issue.
  • UI and Mockup Comparison: Provide multiple design drafts or screenshots and request a detailed comparison of differences.
  • Data Extraction: Extract structured text or numbers from charts, diagrams, receipts, or handwriting.
  • Automated Image Tagging: Generate detailed descriptions and tags for local image assets automatically.
  • Mixed-Source Audits: Evaluate an online image against a local mockup by sending both in a single query.

How Vision Works

  1. Trigger Detection: The agent automatically triggers this skill when you reference image paths, upload attachments, or use phrases like "look at this picture" or "compare these mockups".
  2. Image Conversion: The skill reads local images and converts them into standard base64 strings, or prepares public URLs for remote retrieval.
  3. API Dispatch: The skill dispatches the visual data payload along with your prompt to a configured external vision service (such as Alibaba Cloud's Bailian or OpenAI) using an OpenAI-compatible format.
  4. Context Injection: The external vision API generates a textual description or answers the query, which is then fed back to your primary AI model to proceed with the conversation.

Vision Setup

Configure and initialize the vision service using the command-line helper. This is a standard setup workflow within the Openclaw Skills architecture.

Quick Setup

Run the setup script to input your external vision API key, base URL, and preferred model name:

node scripts/vision.js --setup

Verify Configuration

Inspect your active settings and environment variables anytime with the following command:

node scripts/vision.js --config

Running Commands

Execute manual visual analysis using local or remote paths:

# Analyze a local image
node scripts/vision.js /path/to/image.jpg "Describe the contents"

# Analyze a web-hosted image
node scripts/vision.js --url https://example.com/image.png "What is this?"

# Multi-image comparative analysis
node scripts/vision.js image1.png image2.png "Compare these two images"

Vision Data Schema & Taxonomy

The Vision skill stores and manages configuration locally in a dedicated user directory. Ensure you have Node.js installed on your system.

Configuration Directory

  • Path: ~/.claude/skills/vision/config.json
  • Structure:
{
  "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
  "api_key": "your_api_key_here",
  "model": "qwen3.5-omni-plus"
}

Supported Image Formats

Format Extension Notes
JPEG .jpg, .jpeg Fully supported, optimal compression
PNG .png Fully supported, ideal for screenshots
GIF .gif Supported (static analysis)
WebP .webp Fully supported
BMP .bmp Supported

Compatible Visual APIs

Provider Recommended Model Benefits
Alibaba Cloud Bailian qwen3.5-omni-plus Highly recommended; includes free trial tokens for new users
Alibaba Cloud Bailian qwen-vl-max Outstanding multilingual vision performance
OpenAI gpt-4o-mini Extremely fast, reliable global model (requires overseas billing)
Custom Any OpenAI-compatible model Highly flexible; just update the BASE_URL and model in your config

Vision Advanced Features

  • Multi-Image Comparison Mode: Pass multiple images simultaneously to compare visual elements, layouts, and style changes side-by-side.
  • Hybrid Source Processing: Seamlessly mix local file paths and remote URLs within a single analysis prompt.
  • Flexible Backend Compatibility: Supports any API provider conforming to the standard OpenAI chat completions payload structure, making it highly customizable within the Openclaw Skills framework.
  • Automatic Prompt Context Routing: Automatically detects user intent (such as screenshots or attachments) to invoke the vision skill dynamically without requiring manual script calls.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*