OpenClaw Voice for Openclaw

A CLI-driven voice interface for AI agents utilizing Whisper for transcription and ElevenLabs for high-quality speech synthesis.

frank-bot07
v1.0.0
Feb 20, 2026
2
1.5k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install openclaw-voice

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install openclaw-voice using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is OpenClaw Voice?

OpenClaw Voice is a specialized utility designed to integrate advanced audio capabilities into your terminal-based AI workflows. By leveraging Whisper STT (Speech-to-Text) and ElevenLabs TTS (Text-to-Speech), this skill allows developers to record, transcribe, and synthesize spoken language. It is a vital addition for users building a comprehensive ecosystem of Openclaw Skills that require audio-native interactions.

Architected with a focus on performance and minimal dependencies, the skill bypasses heavy external audio libraries by utilizing system-level tools like sox and ffplay. This ensures that the voice conversation layer remains lightweight while maintaining high fidelity for both transcription accuracy and vocal response quality.

OpenClaw Voice Use Cases

  • Transcribing voice commands or notes into text for AI processing
  • Generating natural-sounding vocal responses from LLM outputs
  • Maintaining a searchable database of voice-to-text transcripts
  • Building voice-enabled automation scripts within the terminal
  • Extending the capabilities of other Openclaw Skills with audio triggers

How OpenClaw Voice Works

  1. The user triggers a recording command via the CLI built with commander.
  2. System tools like sox/rec are called via child_process to capture audio input.
  3. The audio is processed through the Whisper STT engine to generate a text transcript.
  4. Transcripts and associated metadata are stored in a better-sqlite3 database using Write-Ahead Logging (WAL) for speed.
  5. To respond, text is sent to the ElevenLabs API to generate a high-quality audio file.
  6. The resulting audio is played back to the user via ffplay.

OpenClaw Voice Setup

Before installation, ensure you have sox and ffmpeg installed on your host system. Navigate to your skill directory and run the following commands:

npm install
npm run migrate

You will need to configure your environment variables with valid API keys for Whisper and ElevenLabs to enable the transcription and synthesis features within your Openclaw Skills environment.

OpenClaw Voice Data Schema & Taxonomy

The skill organizes its data using a local SQLite database to ensure fast retrieval and local persistence. The schema is optimized for transcript management:

Table Description
transcripts Stores unique UUIDs, the transcribed text, and timestamps.
audio_files Maps text records to local file paths for generated speech.
settings Stores configuration metadata for STT and TTS engines.

OpenClaw Voice Advanced Features

  • SQLite WAL mode for high-concurrency database operations
  • Integration with the @openclaw/interchange protocol for data sharing across Openclaw Skills
  • Lightweight audio handling using native child_process execution
  • Extensible CLI architecture for custom voice-based workflows
  • Planned support for real-time conversation mode in future releases

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*