A specialized tool for understanding and summarizing local PDF, video, and audio files using the Google Gemini API.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install gemini-reader
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install gemini-reader using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
Gemini Reader is a high-performance extension designed for the Openclaw Skills framework, enabling AI agents to process non-text assets that fall outside standard LLM context windows. By utilizing the Gemini Python SDK, this skill allows users to extract intelligence from complex PDF documents, video files, and audio recordings. It serves as a bridge for developers who need to incorporate deep multimodal analysis into their automated workflows without manual data entry.
This skill is particularly valuable for users of Openclaw Skills who require robust transcription and summarization capabilities for media. It handles the heavy lifting of file uploading, processing via Google's latest models like Gemini 2.5 and 3.1, and ensuring temporary files are cleaned up post-analysis for maximum efficiency.
To integrate this capability into your Openclaw Skills environment, ensure you have the required Python SDK and API credentials configured.
# Install the Gemini Python SDK
pip install google-genai
# Set your API key as an environment variable
export GEMINI_API_KEY='your_gemini_api_key_here'
Gemini Reader supports a wide range of media formats and organizes data through a standardized processing pipeline.
| Media Type | Supported Extensions | Common Metadata Handled |
|---|---|---|
| Documents | Page count, text content, embedded tables | |
| Video | .mp4, .webm, .mov, .avi, .mkv | Visual cues, spoken dialogue, timestamps |
| Audio | .mp3, .wav, .m4a, .ogg | Speech-to-text, speaker tone, audio duration |
Output can be directed to the terminal or a persistent text file using the --output flag.
Loading
A powerful, API-free internet search tool for AI agents to retrieve real-time data from Baidu and Bing.

DocStrange is a powerful document extraction API by Nanonets that converts PDFs and images into structured Markdown, JSON, or CSV data with field-level confidence scoring.

A fully local, CPU-based text-to-speech engine providing high-quality audio generation and voice cloning without internet dependencies.

A multi-step deployment agent that orchestrates the workflow from local build and testing to GitHub hosting and Cloudflare Pages deployment.

A structured 5-phase development framework designed to guide AI agents from vague requirements to production-ready code and iteration.

A systematic 5-phase framework designed to transform vague ideas into production-ready software through structured AI collaboration.








































