Video Watcher for Openclaw

An automated utility for downloading videos, generating high-accuracy transcripts with Whisper, and extracting frames for visual analysis.

zedit42
v1.0.0
Feb 24, 2026
0
0
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install xeonen-video-analyzer

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install xeonen-video-analyzer using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Video Watcher?

Video Watcher is a sophisticated media processing tool designed to bridge the gap between video content and textual analysis. By integrating industry-leading utilities such as yt-dlp, ffmpeg, and OpenAI Whisper, it allows developers to programmatically ingest video data from various platforms and convert it into a structured format. This skill is a vital component for anyone building automated research pipelines or documentation engines within the Openclaw Skills ecosystem, providing a reliable way to handle high-volume video processing tasks.

The core strength of Video Watcher lies in its multi-modal approach. It doesn't just download a file; it deconstructs the media into its constituent parts: raw video, isolated audio, timestamped subtitles, and periodic visual snapshots. This granular output makes it significantly easier for AI agents or human researchers to parse long-form content, identify key moments, and generate summaries without manually watching hours of footage.

Video Watcher Use Cases

  • Performing due diligence (DD) on company presentations and technical product demos.
  • Generating comprehensive lecture notes and study materials from educational videos.
  • Transforming podcast recordings into searchable text archives and summaries.
  • Creating technical documentation for software tutorials by capturing both steps and screenshots.
  • Processing internal meeting recordings for automated action-item extraction.

How Video Watcher Works

  1. The user triggers the workflow by providing a video URL to the analysis script.
  2. yt-dlp fetches the media from the source URL and saves it to the local output directory.
  3. ffmpeg extracts the audio track from the video file to ensure optimal transcription quality.
  4. OpenAI Whisper processes the extracted audio to generate both plain text transcripts and SRT subtitle files.
  5. Simultaneously, ffmpeg captures frame screenshots at a defined interval (e.g., every 30 seconds) to provide visual context.
  6. The resulting assets are organized into a standardized folder structure, ready for further AI-assisted summarization or manual review.

Video Watcher Setup

To utilize this entry in the Openclaw Skills collection, you must first install the required binary dependencies on your system:

brew install yt-dlp ffmpeg openai-whisper

After installation, you can initiate a video analysis by running the provided shell script:

./scripts/analyze.sh "https://youtube.com/watch?v=example"

Video Watcher Data Schema & Taxonomy

All processed data is stored in the outputs/ directory with a standardized taxonomy for easy retrieval:

Asset File Name Description
Video video.mp4 The raw downloaded media file.
Audio audio.mp3 The extracted audio stream used for processing.
Text transcript.txt A plain text version of the speech-to-text output.
Subtitles transcript.srt Time-coded subtitles for sync-heavy workflows.
Visuals frames/ A sub-directory containing periodic JPG screenshots.

Video Watcher Advanced Features

  • Support for multiple Whisper model sizes (tiny, base, medium, large) to optimize for speed or accuracy.
  • Adjustable frame-capture intervals via configuration to match the pace of different video types.
  • Direct integration with AI agents for immediate summarization using the clawdbot ask command.
  • Flexible output directory management to support batch processing and organizational workflows.
  • Compatibility with a wide range of video hosting platforms supported by the yt-dlp backend.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*