TokenGuard is a lightweight prevention engine that intercepts LLM API requests to avoid rate limits and optimize token usage efficiency.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install token-guard
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install token-guard using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
TokenGuard serves as a critical middleware layer designed to prevent the common 429 (Rate Limit / Resource Exhausted) errors that plague AI development. By acting as a pre-flight gateway, it allows developers to manage high-performance LLM interactions without falling into "death spiral" retry loops. It is an essential component for those building with Openclaw Skills who require maximum intelligence while minimizing API costs.
The engine is specifically optimized for efficiency, featuring a zero-dependency architecture that fits perfectly into resource-constrained environments. Whether you are using Google Gemini, Anthropic Claude, or OpenAI GPT-4o, TokenGuard provides a localized way to track quotas and handle model fallbacks before a request ever hits the network.
Integrating TokenGuard into your environment is straightforward as it requires no external libraries. Add the skill to your agent configuration to begin using Openclaw Skills for rate limit management.
skills:
- token-guard
In your Python application logic:
from token_guard import TokenGuard
guard = TokenGuard()
decision = guard.check(prompt_text, model="gemini-3-flash")
if decision.action == "proceed":
# Execute your API call here
guard.record_usage(decision.estimated_tokens, model="gemini-3-flash")
TokenGuard maintains a localized schema to track model health and usage statistics. This data allows for precise monitoring of the efficiency gained through Openclaw Skills.
| Field | Type | Description |
|---|---|---|
| tpm_limit | Integer | The maximum tokens permitted per minute for the selected model. |
| used_this_minute | Integer | The count of tokens consumed in the current 60-second window. |
| usage_pct | String | Percentage-based representation of quota consumption. |
| tokens_saved | Integer | The total number of tokens saved through duplicate detection and caching. |
| status | String | Current health status (OK, Warning, or Blocked) of the API connection. |
Loading
A streamlined utility pack designed to install curated bundles of AI agent skills in a single command.

A lightweight Python engine that orchestrates multi-agent squads by routing tasks to the most efficient model based on skill requirements and cost.

A high-efficiency financial tracking engine for AI agents to monitor API costs, gas fees, and revenue in real-time.

A public-safe orchestration tool for managing three-role agent squads with stable pseudonyms and structured JSON outputs.

An interactive Telegram skill that presents multiple clarifying questions as a structured form with inline buttons and free-text support.

An AI-powered medical assistant providing educational insights into symptoms, diseases, medications, and general health maintenance.








































