AI Agent Hub
Back to skills
AI Agent Security Red Teaming icon

AI Agent Security Red Teaming

IT Ops & Security Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow the guide at https://skillhub.cn/install/skillhub.md to install @user_4b84a253/123123123123123123 into your AI assistant.

About this skill

Problem to Solve

When AI Agents can load third-party Skills, the risk often comes from the Skill's prompts, tool parameters, or external links rather than the model alone. A malicious Skill may attempt jailbreaks, bypass safeguards, or direct the Agent toward untrusted URLs. Functional tests may show whether features work, but not whether the Agent can resist abuse.

How the Skill Works

This Skill is oriented toward red-team security testing for AI Agents. It frames checks around jailbreak defense, malicious instruction detection, and URL access control, allowing reviewers to evaluate behavior after third-party Skills are loaded.

  • Threat perspective: examines jailbreak attempts, malicious instructions, and suspicious URL access.
  • Evaluation approach: observes whether the Agent refuses, downgrades, audits, or blocks abnormal requests under controlled inputs.

Scope and Caveats

It focuses on post-Skill security defense and agent behavior triggered by third-party inputs. If the tested system lacks logging, permission isolation, or policy enforcement, results only show whether current controls trigger; they do not replace full penetration testing or guarantee production safety.

Use Cases

  • Before release, send jailbreak samples to an AI Agent to verify it still refuses risky instructions after loading third-party Skills.
  • Build Skill prompts containing malicious instructions to see whether the Agent detects anomalies and blocks tool calls or output.
  • Inject Skill parameters pointing to suspicious external URLs to check whether the Agent enforces URL access controls.
  • Re-test Skill injection scenarios during red-team reviews and record the Agent's refusal, downgrading, or audit behavior.

Best For

  • Platform security engineers responsible for AI Agent security reviews, who need to verify jailbreak defense after third-party Skills are loaded.
  • IT operations security leads managing internal LLM Agents, who need to check malicious instruction detection and URL controls after Skill loading.
  • LLM application testers who need to inject Skill attack samples into Agents and record security responses.
  • Security operations staff maintaining audit workflows, who need to review Agent blocking records for suspicious URLs and abnormal instructions.