Jarvis TTS (Text-to-Speech) for Openclaw

A high-performance text-to-speech engine for macOS utilizing Microsoft edge-tts to deliver human-like neural voice output.

e421083458
v1.0.0
Feb 22, 2026
2
1.5k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install jarvis-tts

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install jarvis-tts using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Jarvis TTS (Text-to-Speech)?

Jarvis TTS is a sophisticated speech synthesis tool specifically optimized for macOS users within the Openclaw Skills ecosystem. By leveraging the advanced Microsoft edge-tts engine, it produces high-fidelity, neural voice outputs that provide a significantly more natural listening experience compared to standard system voices. This skill is designed to give AI agents a realistic voice, enabling them to communicate through audio responses, narrate long-form text, or provide audible system notifications.

The tool is built as a developer-friendly wrapper that handles the complexity of API calls, audio buffering, and playback synchronization. It is particularly effective for users requiring high-quality Chinese language support, offering various personas ranging from professional news anchors to lively, energetic voices. As a key component of Openclaw Skills, it provides a seamless bridge between text-based AI processing and interactive audio feedback.

Jarvis TTS (Text-to-Speech) Use Cases

  • Creating interactive AI assistant voice responses for hands-free operation.
  • Converting written documentation or articles into high-quality audio files for listening on the go.
  • Setting up audible system alerts or reminders that use natural language instead of simple beeps.
  • Developing accessibility features for macOS-based applications and terminal workflows.

How Jarvis TTS (Text-to-Speech) Works

  1. The user inputs text and selects an optional voice profile via the command line or script interface.
  2. The skill communicates with the Microsoft TTS API using the edge-tts library to generate a high-quality MP3 file.
  3. The system performs a file integrity check to ensure the audio was generated correctly and meets size requirements.
  4. The macOS afplay utility is invoked to play the audio file synchronously, ensuring the script waits until the speech is finished.
  5. Temporary audio files are purged from the system after playback to maintain a clean environment.

Jarvis TTS (Text-to-Speech) Setup

To integrate this into your collection of Openclaw Skills, ensure you are on macOS and have Python 3 installed. Follow these steps:

  1. Install the core edge-tts dependency:
pip3 install edge-tts
  1. Make the shell wrapper executable:
chmod +x jarvis-tts.sh
  1. Test the installation with a simple command:
./jarvis-tts.sh "Hello, I am ready to assist you."

Jarvis TTS (Text-to-Speech) Data Schema & Taxonomy

The skill manages data through transient audio files and structured voice parameters. Below is the metadata structure used during execution:

Attribute Type Description
Input Text String The raw text to be converted to speech (UTF-8 support)
Voice ID String The Microsoft Neural voice identifier (e.g., zh-CN-YunxiNeural)
Temp File Path Location of the generated MP3 in /tmp/
Timeout Integer Default 60-second limit for voice generation tasks

Jarvis TTS (Text-to-Speech) Advanced Features

  • Access to a wide range of Microsoft Neural voices, including specialized male (Yunxi, Yunjian, Yunyang) and female (Xiaoxiao, Xiaoyi) profiles.
  • Automatic playback duration calculation to prevent process hanging during long narrations.
  • Seamless integration with other Openclaw Skills for automated voice reporting of script results.
  • Support for offline repetition of generated audio by bypassing the API call if the file is retained.
  • Extension-ready architecture allowing for easy porting to Linux or Windows playback utilities.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*