kube-medic for Openclaw

A powerful Kubernetes diagnostic toolkit for instant AI-driven cluster triage and incident response using standard kubectl commands.

tkuehnl
v1.0.3
Feb 22, 2026
3
2k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install kube-medic

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install kube-medic using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is kube-medic?

kube-medic is a comprehensive Kubernetes diagnostics skill designed to turn AI agents into expert Site Reliability Engineers. It provides a structured interface for performing full cluster health triage, detailed pod autopsies, and resource pressure analysis. By correlating data across events, logs, and deployment histories, it enables fast root-cause identification for complex infrastructure issues.

Integrating seamlessly with Openclaw Skills, it provides a safe yet powerful way to manage production environments with built-in safeguards for write operations. The tool focuses on data correlation, ensuring that symptoms like CrashLoopBackOff are linked to underlying causes like OOMKilled events or misconfigured resource limits, providing a professional-grade diagnostic experience.

kube-medic Use Cases

  • Identifying why specific pods are stuck in CrashLoopBackOff or ImagePullBackOff states.
  • Performing a rapid cluster-wide health sweep to detect NotReady nodes or resource pressure.
  • Analyzing stuck deployments to determine if a recent rollout needs an immediate rollback.
  • Monitoring real-time cluster events to correlate recent changes with system failures.
  • Safely executing administrative actions like restarting deployments or scaling replicas after AI diagnosis.

How kube-medic Works

  1. The agent initiates a cluster-wide sweep to identify critical issues like failing pods, NotReady nodes, or recent warning events.
  2. If specific problems are found, the tool performs a deep-dive autopsy on affected pods, correlating current logs, previous logs, and container states.
  3. The system analyzes resource consumption (CPU/Memory) across the cluster to determine if hardware limits or node pressure are impacting performance.
  4. The AI correlates data from multiple subcommands to provide a coherent diagnosis and actionable fix rather than just listing raw symptoms.
  5. The user reviews the proposed resolution and provides explicit confirmation before any write commands, such as rollbacks or restarts, are executed.

kube-medic Setup

To utilize this within the Openclaw Skills ecosystem, ensure you have the necessary CLI dependencies installed and your kubeconfig is properly configured.

# Install required dependencies
sudo apt-get install kubectl jq

# Verify cluster connectivity
kubectl cluster-info

# The skill executes via the provided bash script
bash scripts/kube-medic.sh sweep

kube-medic Data Schema & Taxonomy

The skill returns structured JSON data which is then parsed into actionable reports. The following data points are prioritized:

Data Category Key Information Captured
Cluster Sweep Node health, problem pod counts, and high-priority warning events.
Pod Autopsy Container statuses, log tails (current/previous), and image version mismatches.
Deployment Rollout status, generation vs observed generation, and replica set history.
Resource Metrics Node-level CPU/Memory pressure and top 20 resource-consuming pods.
Event Timeline Chronological event logs with summary statistics and top reason counts.

kube-medic Advanced Features

  • Multi-cluster support allowing seamless context switching across different environments via the --context flag.
  • Discord v2 delivery mode optimization for compact triage summaries and interactive quick-action components.
  • Safety-first write operations requiring explicit human-in-the-loop approval before executing rollbacks or pod deletions.
  • Intelligent context management for large-scale clusters to prevent output saturation and focus on high-impact failures.
  • Comprehensive RBAC error handling that provides specific guidance on missing permissions for SRE workflows.

SKILL.md


Loading

Related Openclaw Skills

Featured*