Faster Whisper for Openclaw

A high-performance CTranslate2 reimplementation of OpenAI Whisper providing ultra-fast local transcription, speaker diarization, and advanced subtitle generation.

theplasmak
v1.5.1
Feb 19, 2026
5
8.6k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install faster-whisper

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install faster-whisper using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Faster Whisper?

Faster Whisper is a powerful local speech-to-text engine that utilizes CTranslate2 to achieve 4-6x the speed of standard OpenAI Whisper while maintaining identical accuracy levels. As a premier tool within the Openclaw Skills collection, it provides developers and creators with a robust, cost-effective alternative to cloud-based APIs. By leveraging GPU acceleration (CUDA), the skill can reach up to 20x realtime transcription speeds, meaning a 10-minute audio file can be processed in approximately 30 seconds.

This skill is designed for privacy-conscious environments and high-volume workflows. It supports over 99 languages, automatic language detection, and features advanced processing capabilities like speaker identification, filler word removal, and direct subtitle burn-in. Whether you are managing podcast archives or generating real-time meeting transcripts, this tool delivers professional-grade results locally.

Faster Whisper Use Cases

  • Rapidly transcribing meetings, interviews, and lectures with high accuracy.
  • Generating broadcast-standard subtitles in formats like SRT, VTT, and TTML.
  • Automating podcast transcription by fetching audio directly from RSS feeds.
  • Identifying different speakers in a multi-person conversation using diarization.
  • Translating international audio content directly into English text.
  • Batch processing entire directories of audio or video files for archival search.

How Faster Whisper Works

  1. The skill ingests audio from local files, directories, or remote URLs via yt-dlp.
  2. It applies Voice Activity Detection (VAD) to intelligently identify and isolate speech segments from silence.
  3. The CTranslate2 backend executes the transcription using optimized Whisper models (defaulting to distil-large-v3.5).
  4. If enabled, a speaker diarization pass identifies unique voices, and wav2vec2 alignment refines word-level timestamps to 10ms precision.
  5. The engine formats the output into the requested schema (Text, JSON, CSV, etc.) and saves it to the specified destination.

Faster Whisper Setup

Install the skill and its core dependencies using the provided setup script:

./setup.sh

For advanced speaker diarization capabilities, use the diarize flag during installation:

./setup.sh --diarize

Note: A CUDA-compatible GPU is highly recommended for maximum performance. Use ./setup.sh --check to verify your system compatibility.

Faster Whisper Data Schema & Taxonomy

Component Description
text The full, unstructured transcript of the audio content.
segments A detailed list of objects including start/end timestamps and confidence scores.
language The detected or manually specified ISO language code.
speakers Array of speaker labels (e.g., SPEAKER_1) mapped to their specific dialogue.
stats Metadata regarding processing duration, realtime factor (RTF), and word counts.

Faster Whisper Advanced Features

  • Multi-format concurrent output (e.g., generate SRT and TXT in a single pass).
  • Intelligent filler word removal to strip 'ums' and 'uhs' from the final text.
  • Automatic chapter detection based on configurable silence gaps for long-form content.
  • Stereo channel selection to isolate specific speakers in dual-track recordings.
  • Subtitle line-wrapping and character limits for professional video production standards.
  • Fuzzy search capabilities to find specific terms and timestamps within long transcripts.

SKILL.md


Loading

Related Openclaw Skills

Featured*