AI Agent Security Red Teaming
Paste the following prompt into your AI chat to install this skill:
Please follow the guide at https://skillhub.cn/install/skillhub.md to install @user_4b84a253/123123123123123123 into your AI assistant.
About this skill
Problem to Solve
When AI Agents can load third-party Skills, the risk often comes from the Skill's prompts, tool parameters, or external links rather than the model alone. A malicious Skill may attempt jailbreaks, bypass safeguards, or direct the Agent toward untrusted URLs. Functional tests may show whether features work, but not whether the Agent can resist abuse.
How the Skill Works
This Skill is oriented toward red-team security testing for AI Agents. It frames checks around jailbreak defense, malicious instruction detection, and URL access control, allowing reviewers to evaluate behavior after third-party Skills are loaded.
- Threat perspective: examines jailbreak attempts, malicious instructions, and suspicious URL access.
- Evaluation approach: observes whether the Agent refuses, downgrades, audits, or blocks abnormal requests under controlled inputs.
Scope and Caveats
It focuses on post-Skill security defense and agent behavior triggered by third-party inputs. If the tested system lacks logging, permission isolation, or policy enforcement, results only show whether current controls trigger; they do not replace full penetration testing or guarantee production safety.
Use Cases
- Before release, send jailbreak samples to an AI Agent to verify it still refuses risky instructions after loading third-party Skills.
- Build Skill prompts containing malicious instructions to see whether the Agent detects anomalies and blocks tool calls or output.
- Inject Skill parameters pointing to suspicious external URLs to check whether the Agent enforces URL access controls.
- Re-test Skill injection scenarios during red-team reviews and record the Agent's refusal, downgrading, or audit behavior.
Best For
- Platform security engineers responsible for AI Agent security reviews, who need to verify jailbreak defense after third-party Skills are loaded.
- IT operations security leads managing internal LLM Agents, who need to check malicious instruction detection and URL controls after Skill loading.
- LLM application testers who need to inject Skill attack samples into Agents and record security responses.
- Security operations staff maintaining audit workflows, who need to review Agent blocking records for suspicious URLs and abnormal instructions.
Related Skills
Detects child climbing, leaning out, or gripping window/balcony edges from surveillance video and outputs tiered alerts with historical reports.
A pre-release security auditor for Skills that statically checks injection, credentials, SSRF, CVEs, and permissions, with scored reports.
For independent developers, automates Git weekly reports, prioritized bug tickets, and project health checks into shareable Markdown.
Scan Windows caches, temporary files, and junk files, show space usage and risk levels, and clean selected items to free disk space.