A high-performance text-to-speech engine for macOS utilizing Microsoft edge-tts to deliver human-like neural voice output.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install jarvis-tts
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install jarvis-tts using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
Jarvis TTS is a sophisticated speech synthesis tool specifically optimized for macOS users within the Openclaw Skills ecosystem. By leveraging the advanced Microsoft edge-tts engine, it produces high-fidelity, neural voice outputs that provide a significantly more natural listening experience compared to standard system voices. This skill is designed to give AI agents a realistic voice, enabling them to communicate through audio responses, narrate long-form text, or provide audible system notifications.
The tool is built as a developer-friendly wrapper that handles the complexity of API calls, audio buffering, and playback synchronization. It is particularly effective for users requiring high-quality Chinese language support, offering various personas ranging from professional news anchors to lively, energetic voices. As a key component of Openclaw Skills, it provides a seamless bridge between text-based AI processing and interactive audio feedback.
To integrate this into your collection of Openclaw Skills, ensure you are on macOS and have Python 3 installed. Follow these steps:
pip3 install edge-tts
chmod +x jarvis-tts.sh
./jarvis-tts.sh "Hello, I am ready to assist you."
The skill manages data through transient audio files and structured voice parameters. Below is the metadata structure used during execution:
| Attribute | Type | Description |
|---|---|---|
| Input Text | String | The raw text to be converted to speech (UTF-8 support) |
| Voice ID | String | The Microsoft Neural voice identifier (e.g., zh-CN-YunxiNeural) |
| Temp File | Path | Location of the generated MP3 in /tmp/ |
| Timeout | Integer | Default 60-second limit for voice generation tasks |
Loading
A terminal-based automation tool that uses Playwright to search Bilibili and launch video links directly in your default macOS browser.

A multi-agent AI team designed to automate industry monitoring, content creation, and trend forecasting with user-adaptive intelligence.

A shared 3D virtual environment enabling AI agents to interact, chat, and participate in a collaborative crafting ecosystem as lobster avatars.

Identify songs from humming, singing, or recorded audio files using advanced iFlytek ACRCloud technology.

An efficient AI agent skill that transforms raw notes into professional daily reports, weekly updates, and meeting summaries.

An autonomous task execution system that transitions AI agents from passive heartbeat monitoring to proactive, queue-based workflow management.








































