TokenGuard for Openclaw

TokenGuard is a lightweight prevention engine that intercepts LLM API requests to avoid rate limits and optimize token usage efficiency.

edmonddantesj
v1.5.0
Feb 13, 2026
0
2k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install token-guard

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install token-guard using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is TokenGuard?

TokenGuard serves as a critical middleware layer designed to prevent the common 429 (Rate Limit / Resource Exhausted) errors that plague AI development. By acting as a pre-flight gateway, it allows developers to manage high-performance LLM interactions without falling into "death spiral" retry loops. It is an essential component for those building with Openclaw Skills who require maximum intelligence while minimizing API costs.

The engine is specifically optimized for efficiency, featuring a zero-dependency architecture that fits perfectly into resource-constrained environments. Whether you are using Google Gemini, Anthropic Claude, or OpenAI GPT-4o, TokenGuard provides a localized way to track quotas and handle model fallbacks before a request ever hits the network.

TokenGuard Use Cases

  • Processing large docx or PDF files that risk exceeding per-minute token (TPM) limits.
  • Managing multiple LLM agents that could inadvertently trigger runaway request loops.
  • Reducing costs for developers on free or low-tier API plans by caching identical requests.
  • Ensuring high availability by automatically switching to fallback models when primary quotas are exhausted.
  • Extracting and respecting precise retry delays from API error responses to synchronize request timing.

How TokenGuard Works

  1. The engine performs a pre-flight token estimation using CJK-aware logic to predict usage before the API call is made.
  2. It checks the estimated tokens against a real-time sliding window quota tracker for the specific model being used.
  3. Based on the current usage percentage, it issues a decision: proceed, wait (throttle), fallback to a cheaper model, or block the request.
  4. If the request proceeds, the usage is recorded and the successful response is cached to satisfy future duplicate queries.
  5. In the event of a 429 error from the provider, the engine parses the response to adjust internal timers and prevent further failures.

TokenGuard Setup

Integrating TokenGuard into your environment is straightforward as it requires no external libraries. Add the skill to your agent configuration to begin using Openclaw Skills for rate limit management.

skills:
  - token-guard

In your Python application logic:

from token_guard import TokenGuard

guard = TokenGuard()
decision = guard.check(prompt_text, model="gemini-3-flash")

if decision.action == "proceed":
    # Execute your API call here
    guard.record_usage(decision.estimated_tokens, model="gemini-3-flash")

TokenGuard Data Schema & Taxonomy

TokenGuard maintains a localized schema to track model health and usage statistics. This data allows for precise monitoring of the efficiency gained through Openclaw Skills.

Field Type Description
tpm_limit Integer The maximum tokens permitted per minute for the selected model.
used_this_minute Integer The count of tokens consumed in the current 60-second window.
usage_pct String Percentage-based representation of quota consumption.
tokens_saved Integer The total number of tokens saved through duplicate detection and caching.
status String Current health status (OK, Warning, or Blocked) of the API connection.

TokenGuard Advanced Features

  • Smart Throttle logic that automatically pauses requests when usage exceeds 80% to avoid hard blocks.
  • Multi-model support with pre-configured quotas for Gemini, Claude, GPT-4o, and DeepSeek.
  • CJK-aware estimation that provides accurate token counts for Asian languages without requiring heavy dependencies like tiktoken.
  • Duplicate detection system that identifies and blocks runaway loops (3+ identical requests within 60 seconds).
  • Response caching for successful requests to eliminate redundant token expenditure.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*