Automate the creation of high-quality talking head videos featuring AI avatars, synchronized lipsync, and professional voiceovers.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install talking-head-production
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install talking-head-production using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
Talking Head Production is a comprehensive toolset within the Openclaw Skills ecosystem designed to streamline the creation of virtual presenters and AI spokespeople. By leveraging state-of-the-art models like OmniHuman and PixVerse, this skill allows developers and content creators to transform static portraits and text scripts into dynamic, life-like video content with precise lip-syncing and natural gestures.
Whether you are building course content, social media snippets, or professional product demos, these Openclaw Skills provide the necessary infrastructure to handle complex video synthesis via a simple command-line interface. It integrates seamlessly with text-to-speech engines and video post-processing tools to deliver production-ready results without the need for manual editing suites.
Install the necessary CLI tools and authenticate your session to begin using these Openclaw Skills:
curl -fsSL https://cli.inference.sh | sh && infsh login
After installation, you can verify your setup by running a test dialogue generation and synthesis command to ensure all model dependencies are accessible.
The skill processes various media assets and metadata to produce final video files. Data is organized as follows:
| Asset Type | Description | Key Specifications |
|---|---|---|
| Portrait | Source image for the avatar | Min 512x512, 1024x1024+ recommended, PNG/JPG |
| Audio | Voiceover or dialogue file | 44.1kHz/48kHz, MP3 or WAV format |
| Video Output | Final synthesized talking head | MP4 format, typically 30s segments |
| Metadata | Configuration and model IDs | JSON input for CLI tools |
Loading
An AI-driven workflow for generating cinematic storyboards and shot lists using CLI tools.

A high-performance speech-to-text skill leveraging Whisper models for fast, multi-language audio transcription and translation via the command line.

A developer-centric skill for generating high-conversion multi-slide social media carousels using HTML-to-image automation and AI-driven design rules.

A data-driven automation skill for creating comprehensive SEO content briefs using real-time SERP analysis and keyword research.

A comprehensive text-to-speech skill for AI agents to generate natural, expressive, and multi-speaker audio via the inference.sh platform.

A professional collection of React and Next.js components designed to visualize the complete lifecycle of AI agent tool executions.








































