YouTube Long Video Transcript & Translation for Openclaw

A specialized workflow for extracting, translating, and summarizing verbatim transcripts from long YouTube videos using automated agent orchestration.

qingliu1617-art
v1.0.0
Feb 10, 2026
0
2.1k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install ytb-transcript-long

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install ytb-transcript-long using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is YouTube Long Video Transcript & Translation?

This skill provides a robust framework for handling extended video content that typically exceeds standard session limits. By integrating with the DownSub API, it allows users to pull accurate English subtitles and convert them into high-quality Chinese translations. It leverages sub-agent spawning to manage massive datasets, ensuring that videos over an hour long are processed without timeouts or context loss.

Built as one of the more sophisticated Openclaw Skills, this workflow prioritizes verbatim accuracy while adding value through automated summarization. It is designed to bridge the gap between raw video data and structured, readable documentation, making it an essential tool for researchers and content consumers who need to localize long-form media.

YouTube Long Video Transcript & Translation Use Cases

  • Extracting full verbatim subtitles from long-form educational or technical YouTube content.
  • Translating English transcripts into Chinese for localized documentation or research purposes.
  • Generating executive summaries and key metric tables from extensive video data automatically.
  • Streamlining the workflow of subtitle extraction and cloud-based document generation for high-volume tasks.

How YouTube Long Video Transcript & Translation Works

  1. Performs an initial environment check for DownSub API authorization and local tool availability.
  2. Verifies the video language via the API to ensure English source tracks are used exclusively for translation.
  3. Assesses the total line count of the transcript to determine if a sub-agent is required for background processing.
  4. Spawns a specialized sub-agent to slice the transcript into manageable 500-line segments.
  5. Executes verbatim translation and formatting of each segment into structured Markdown headers.
  6. Synthesizes metadata including a Chinese executive summary and a key metrics table based on the first 500 lines.
  7. Merges all segments and delivers the final output via the zhiyan platform or a local Markdown file.

YouTube Long Video Transcript & Translation Setup

To configure the DownSub API for this skill, ensure you have your Bearer token ready. This is a primary requirement for Openclaw Skills that interact with external video data.

# Example of how the API key is passed in headers
Authorization: Bearer YOUR_AIza_TOKEN_HERE
Content-Type: application/json

If you wish to generate online documents, ensure the zhiyan MCP is installed and configured in your environment.

YouTube Long Video Transcript & Translation Data Schema & Taxonomy

The skill organizes data into a structured Markdown output with specific metadata sections:

Section Description
Executive Summary A 3-5 bullet point summary of the video content in Chinese.
Key Metrics Table A Markdown table highlighting figures like revenue, growth, or KPIs found in the text.
Full Transcript Chunked, verbatim Chinese translation of the subtitles with timestamp-based headers.
File Taxonomy Output saved as full_transcript.md or parsed via zhiyan for cloud hosting.

YouTube Long Video Transcript & Translation Advanced Features

  • Sub-agent spawning for background processing of videos exceeding 1000 lines to maintain session stability.
  • Automated chunking logic (500-line segments) to prevent LLM context overflows during translation.
  • Conditional routing between local file output and zhiyan cloud document hosting based on tool availability.
  • Integrated metadata extraction for identifying key business metrics within video scripts during the translation phase.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*