Voice Chat Bridge for Openclaw

A robust bi-directional voice conversation system that adds speech recognition and natural synthesis to AI coding agents.

patrickgeek
v1.1.0
Mar 3, 2026
0
1.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install voice-chat-bridge

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install voice-chat-bridge using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Voice Chat Bridge?

The Voice Chat Bridge is a sophisticated integration designed for Openclaw Skills that transforms static text interactions into dynamic voice conversations. By leveraging Edge TTS for high-quality synthesis and tools like hear or DashScope for transcription, it creates a seamless audio loop for AI assistants. Whether you are building a personal companion or a collaborative tool, this skill ensures your agent can hear, understand, and speak back with natural human-like voices.

This bridge is highly versatile, supporting multiple deployment scenarios ranging from offline local playback to globally accessible public endpoints using Cloudflare Tunnels. By integrating this into your workflow, you can move beyond keyboard-based interaction and engage with your AI using natural speech across various platforms including Telegram, Discord, and Slack.

Voice Chat Bridge Use Cases

  • Hands-free interaction with AI agents during intensive coding sessions.
  • Developing multi-platform voice bots for Telegram, Discord, or Slack using Openclaw Skills.
  • Building a voice-enabled home assistant accessible via a local web interface.
  • Creating continuous voice chat loops for natural language practice and interaction.
  • Automating voice-based summaries of daily logs or diaries.

How Voice Chat Bridge Works

  1. Capture: The system listens for voice input via a configurable hotkey (Ctrl+T) or receives incoming audio files from messaging platforms.
  2. Transcription: Audio is processed using ffmpeg and converted to text via local STT tools like hear or cloud-based services.
  3. Intelligence: The transcribed text is sent to the AI agent for processing, context analysis, and response generation.
  4. Synthesis: The AI's text response is converted back into high-quality speech using the Edge TTS engine, supporting over 100 languages.
  5. Delivery: The generated audio is either played locally on the host machine, served via a web UI, or sent back as a voice message to the user.

Voice Chat Bridge Setup

To get started with this among your Openclaw Skills, first install the core system dependencies and Python libraries.

# Install audio processing tools
brew install ffmpeg

# Install speech synthesis library
pip3 install edge-tts

Initialize the skill structure and configuration files:

bash skills/voice-chat-bridge/scripts/init.sh

Finally, configure your voice_config.json in the workspace to set your preferred mode (local, web, or cloud) and choose your voice character ID.

Voice Chat Bridge Data Schema & Taxonomy

The skill manages its state and outputs within the Openclaw workspace directory using a structured approach:

File/Directory Description
voice_config.json Primary configuration for TTS engine, hotkeys, and networking protocols.
voice_output/ Directory containing generated .mp3 and .ogg audio files for playback.
.voice_trigger Internal bridge file used to trigger AI processing from voice input.
habits.json Logs interaction frequency and emotional connection metadata for the agent.

Voice Chat Bridge Advanced Features

  • Global Hotkey Support: System-wide Ctrl+T shortcut for instant push-to-talk functionality regardless of the active window.
  • Emotional Synthesis Integration: Ability to adjust voice style and pitch based on the agent's internal emotional state.
  • Global Tunneling: Native support for Cloudflare Tunnel and Ngrok to provide secure public URLs for audio files.
  • Local Web Interface: Built-in mini-server with a web UI and QR code for easy access from mobile devices on the same network.
  • Multi-language Detection: Automatic switching between 100+ regional accents and languages to match user input.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*