AI Agent Hub
Back to skills
TDD-Driven Agent Skill Writing icon

TDD-Driven Agent Skill Writing

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_3ee04ca3/test-new3 by following https://skillhub.cn/install/skillhub.md.

About this skill

What Problem It Solves

Agent skill documentation often looks clear to a human, but agents may still find plausible excuses and bypass rules under pressure. One-off narratives, multi-language examples, and flowcharts used for linear steps can also make skills harder to retrieve and harder to follow. test2 treats skill authoring as TDD: run a no-skill baseline first, observe the exact failure modes, write only the documentation needed to address those failures, and then refactor against newly discovered loopholes.

How It Works

  • RED: Use a subagent to run multiple combined pressure scenarios, such as time pressure, sunk cost, authority, and fatigue; record the agent’s verbatim rationalizations, choices, and triggers.
  • GREEN: Write only enough guidance to cover the observed failures; avoid speculative content. SKILL.md should have a clear description, trigger conditions, core principle, and actionable constraints.
  • REFACTOR: Every time an excuse appears, such as “I’m following the spirit” or “this is just a simple addition,” add explicit counters, red-flag checks, and CSO symptom language.
  • Type-aware testing: Discipline skills need maximum-pressure tests, technique skills need application and edge cases, pattern skills need recognition and counter-examples, and reference skills need retrieval and correct-usage tests.

Boundaries

This fits reusable agent skills for Claude Code, Codex, or similar agents, especially TDD, verification, debugging, and process constraints. It is not intended for one-off project conventions, mechanical rules that can be automated with regex or validation, or batch skill creation without baseline tests. Prefer one high-quality example over multiple diluted ones; heavy references or reusable scripts can be split into separate files, but flowcharts should not use labels without semantic meaning.

Use Cases

  • Before creating a discipline skill, run no-skill subagent pressure tests and record the agent's exact rationalizations.
  • When editing an existing SKILL.md, rerun the original pressure scenarios and add explicit counters for new loopholes.
  • When authoring a debugging reference skill, design retrieval scenarios and correct-application test cases for the reference type.
  • When validating a root-cause skill, run the same application scenario to verify the agent can follow the minimal document.

Best For

  • AI engineering leads maintaining Claude Code custom rules: need to turn TDD-style discipline into agent-followable skills.
  • SREs writing team process documents for Codex: need debugging, verification, and pre-commit checks to hold under pressure.
  • Developer platform engineers maintaining agent skill libraries: need technique skills split into testable, searchable, reusable files.
  • Ops engineers writing debugging references for internal assistants: need to verify retrieval accuracy and real-scenario application.