TokenRanger for Openclaw

TokenRanger is a local context compression plugin for Openclaw Skills that slashes cloud LLM token costs by up to 80% using local SLMs.

synchronic1
v1.0.0
Mar 1, 2026
0
868
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install tokenranger

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install tokenranger using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is TokenRanger?

TokenRanger is a powerful performance-oriented plugin designed for the Openclaw ecosystem. It acts as a middleware between the user and cloud LLMs, intercepting conversation history and compressing it through a local Ollama instance before transmission. This significantly reduces input token usage while maintaining semantic accuracy, making it one of the most cost-effective Openclaw Skills for heavy developers.

By leveraging Small Language Models (SLMs) like Mistral or Phi-3 locally, TokenRanger ensures that your cloud bills remain low without sacrificing the intelligence of high-end models. It features graceful degradation, meaning if your local service is unavailable, it simply falls back to standard passthrough mode to ensure uninterrupted service.

TokenRanger Use Cases

  • High-volume developers looking to slash API costs for GPT-4 or Claude 3.5 Sonnet
  • Users working with long conversation histories that exceed standard context windows
  • Privacy-conscious environments requiring local pre-processing of session data
  • Systems operating on limited bandwidth where smaller prompt sizes are beneficial

How TokenRanger Works

  1. The OpenClaw gateway receives a user message and triggers the TokenRanger hook.
  2. For the first turn, the full context is sent to ensure high fidelity.
  3. From the second turn onwards, conversation history is sent to a local FastAPI sidecar.
  4. The sidecar utilizes Ollama to run a compression chain via LangChain LCEL.
  5. A semantically dense summary is generated and prepended to the current prompt.
  6. The cloud LLM receives the compressed context, drastically reducing the billable token count.

TokenRanger Setup

To integrate this into your library of Openclaw Skills, follow these steps:

Install the plugin:

openclaw plugins install openclaw-plugin-tokenranger

Run the initial setup to configure the sidecar and pull local models:

openclaw tokenranger setup

Restart your gateway to apply changes:

openclaw gateway restart

Verify the installation:

openclaw tokenranger

TokenRanger Data Schema & Taxonomy

TokenRanger manages its environment and configuration through several key files and services:

Component Path / Detail
Sidecar Service systemd (Linux) or launchd (macOS)
Python Venv ~/.openclaw/extensions/tokenranger/venv
Config File ~/.openclaw/config.json
Local Models mistral:7b (GPU) or phi3.5:3b (CPU)
Sidecar Logs ~/.openclaw/extensions/tokenranger/service/logs

TokenRanger Advanced Features

  • Multi-Strategy Inference: Automatically switches between GPU-based semantic summarization and CPU-based extractive bullet points.
  • Dynamic Context Control: Configure minimum prompt lengths via minPromptLength to prevent compression on short messages.
  • In-Chat Commands: Real-time mode switching and status checks using /tokenranger commands directly in the interface.
  • Robust Fallback: Silent failover to uncompressed context if the local Ollama service becomes unreachable, ensuring reliability.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*