WeVoiceReply for Openclaw

An automated audio pipeline that transforms AI text into natural voice messages and pushes them directly to users via media protocols.

zhairen
v1.0.3
Feb 16, 2026
0
761
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install wevoicereply

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install wevoicereply using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is WeVoiceReply?

WeVoiceReply is a specialized automation tool designed for the Openclaw Skills ecosystem to provide high-quality voice synthesis and delivery. It streamlines the process of converting AI-generated text into audible speech using a robust pipeline involving Piper TTS and FFmpeg. By handling the complexities of audio transcoding (converting WAV to AMR) and automated URL generation, it allows developers to focus on the conversational experience while the skill manages the technical delivery of voice assets.

This skill is particularly effective for creating warm, human-like interactions within any environment using Openclaw Skills. It encourages the use of natural phrasing and specific punctuation to ensure that synthesized audio has realistic pauses and prosody. As part of a larger multi-agent setup, WeVoiceReply acts as the final vocal interface of the system, ensuring that communications are not just read, but heard through a seamless, three-step automated workflow.

WeVoiceReply Use Cases

  • Developing AI agents that need to respond with voice notes rather than standard text.
  • Implementing 'read aloud' features for complex summaries, research findings, or messages.
  • Automating status updates in communication channels where voice is preferred over text for accessibility.
  • Creating professional notification systems requiring high-quality voice-over updates instead of simple alerts.

How WeVoiceReply Works

  1. The AI agent generates a natural, oral-focused text script based on user context, incorporating specific punctuation for delivery.
  2. The skill invokes the internal voice_reply_skill.py script, passing the generated text as a parameter.
  3. Piper TTS synthesizes the text into a raw WAV audio file.
  4. FFmpeg transcodes the WAV file into an AMR format, optimizing it for mobile and messaging protocols.
  5. The file is uploaded to a cloud storage endpoint, returning a JSON object containing the media URL.
  6. A secondary call is automatically made to the system message API to deliver the voice file to the target audience without further manual input.

WeVoiceReply Setup

To use this skill within your Openclaw Skills environment, ensure you are running on a Linux system with the following requirements installed:

# Install system dependencies
sudo apt-get update && sudo apt-get install ffmpeg python3

# Ensure your Python environment has access to Piper
pip install piper-tts

Verify that your FFmpeg installation supports the AMR codec for successful audio transcoding and that your Python path is correctly configured for the skill scripts.

WeVoiceReply Data Schema & Taxonomy

WeVoiceReply organizes its data flow through a structured parameter system and metadata taxonomy:

Parameter Requirement Description
text Mandatory The source text string that will be converted to speech.
url Output The generated link to the processed .amr audio file.
action Internal Logic flag set to 'send' for the final message delivery phase.
contentType Internal Defined as 'voice' to ensure the client renders the audio player correctly.

WeVoiceReply Advanced Features

  • Decoupled architecture separating high-level workflow orchestration from low-level TTS processing logic.
  • Intelligent prosody handling using specific linguistic punctuation (like Chinese commas) for natural speech flow.
  • Automated three-step pipeline (Synthesis -> Transcoding -> Delivery) that prevents manual errors in media handling.
  • Shell-safe parameter passing using single-quote wrapping to handle complex text inputs and special characters.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*