AI Agent Hub
Back to skills
OpenClaw Resilience Monitor icon

OpenClaw Resilience Monitor

IT Ops & Security Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_15292d5a/yjkj-resilience-monitor using https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When agents call LLM APIs, errors such as 429, 503, timeouts, and network failures can interrupt tasks, while raw logs remain scattered across requests and instances. Troubleshooting often requires reading logs, checking per-model behavior, and confirming retry records, making it hard to tell quickly whether a model is degrading, whether an error type is retryable, or whether a failed task was recovered.

How It Works

This skill adds natural-language controls for OpenClaw to inspect API health and manage retry behavior:
- resilience_stats: query today, hour, week, or per-model error statistics.
- resilience_strategies: list, add, update, or reset retry strategies with fixed, exponential, or custom intervals.
- resilience_report: generate daily, per-model, recovery, or full status reports.
- resilience_dashboard: open the local real-time panel for current errors, model breakdowns, recent failures, and strategy cards.
For example, an operator can ask for a model's daily error rate and switch the default timeout strategy to exponential backoff. The skill classifies errors such as rate_limit, timeout, auth_failed, and context_too_long, and marks which categories are retryable. Tool registration, automatic error recording, the retry engine, and the Dashboard service are provided by the companion plugin.

Boundaries

It is intended for engineering workflows already using OpenClaw where API health should be checked with plain Chinese instructions. Without the companion plugin installed and loaded, tool calls will fail and the dashboard will not start. Strategy edits usually apply only to the local Gateway instance.

Use Cases

  • When an OpenClaw gateway shows 429s or timeouts across models, check daily, hourly, or per-model error rates to isolate the issue.
  • During agent retry troubleshooting, inspect fixed, exponential, or custom strategies and adjust timeout retry behavior.
  • Produce a daily or per-model error report explaining whether rate-limit, timeout, and similar failures are retryable or recovered.
  • Open, stop, or check the local OpenClaw dashboard to track active retries and model-level error distribution.

Best For

  • SREs maintaining OpenClaw gateways who need natural-language views of API health, error classes, and retry strategy status.
  • Backend engineers running multi-model agents who must debug 429, timeout, and model-unavailable errors per model.
  • Engineers responsible for agent task reliability who want daily error reports and confirmation that retries or recovery ran.
  • Platform engineers building local AI ops dashboards who need to open the panel and track live errors and strategy cards.