MenuVision for Openclaw

MenuVision transforms restaurant menus from any source into interactive, visually rich HTML digital menus using advanced AI vision and image generation.

ademczuk
v1.0.1
Feb 23, 2026
0
1.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install menuvision

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install menuvision using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is MenuVision?

MenuVision is a comprehensive technical pipeline designed to bridge the gap between physical restaurant menus and modern digital experiences. By leveraging Gemini Vision for data extraction and AI image generation for food photography, it automates the creation of professional HTML menus. This skill is part of the Openclaw Skills collection, providing a robust solution for developers to convert unstructured URLs, PDFs, or photographs into structured JSON data and subsequently into a polished, responsive web interface.

The tool is specifically engineered for high-fidelity extraction, handling complex multi-lingual menus (including CJK scripts) and diverse price formats. It produces a feature-rich output that includes Instagram-style grids, interactive selection receipts, and client-side currency conversion, making it an ideal choice for digital signage, online ordering previews, or restaurant digitization projects.

MenuVision Use Cases

  • Automating the conversion of static PDF menus into interactive digital web pages.
  • Generating high-quality food photography for restaurants with text-only menus.
  • Building multi-lingual digital menus with automatic font fallback for global audiences.
  • Creating self-contained, portable HTML menus for email distribution or offline tablet use.
  • Rapidly digitizing physical restaurant menus from smartphone photographs.

How MenuVision Works

  1. Extraction: The system analyzes the source (URL, PDF, or Photo) using Gemini Vision to generate a structured JSON data contract.
  2. Normalization: The AI agent cleans the data, ensures unique item codes, and normalizes price formats for consistency.
  3. Image Synthesis: For items categorized as food, the pipeline uses Gemini 2.5 Flash Image to generate realistic, casual-style food photography based on item descriptions.
  4. HTML Generation: A build script compiles the JSON data and generated images into a responsive HTML5 file with inline CSS and interactive JavaScript logic.
  5. Deployment: The final menu can be generated as a single portable file or deployed to static hosting platforms like GitHub Pages.

MenuVision Setup

To get started with this skill from the Openclaw Skills library, ensure you have Python 3.9+ installed and configure your environment:

# Install core dependencies
pip install google-genai Pillow requests beautifulsoup4 PyMuPDF

# Install browser automation for JS-heavy sites
pip install playwright && playwright install chromium

# Configure your Google API Key
export GOOGLE_API_KEY="your_api_key_here"

MenuVision Data Schema & Taxonomy

The skill operates on a strict JSON data contract to ensure pipeline reliability. Below is the primary structure:

Component Key Fields Purpose
Restaurant name, cuisine, tagline Defines header branding and image generation context.
Sections title, category, items Groups menu entries; category drives image generation logic.
Items code, name, price, dietary Individual menu entries with pricing and allergen metadata.
Allergen Legend code: display_name Maps technical codes to human-readable allergen names.
Metadata languages, currency Controls CJK font loading and currency display symbols.

MenuVision Advanced Features

  • Portable Mode: Embeds all generated food images as base64 data URIs for a zero-dependency, single-file HTML output.
  • CJK Script Optimization: Automatic detection of Chinese, Japanese, and Korean characters to trigger script-safe image prompting and specific Google Font loading.
  • Client-Side Currency Converter: A built-in toggle in the generated HTML that converts all menu prices using snapshot exchange rates embedded at build time.
  • Intelligent Scraping: Automatically switches between static HTML parsing and Playwright-based screenshots based on website rendering density.
  • Robust API Handling: Includes exponential backoff retry logic and defensive JSON parsing to handle potential LLM output quirks.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Requires
Bins python3
Github Stars: 0
forks: 0

Featured*