AI Agent Hub
Back to skills
Skill Function Test Suite icon

Skill Function Test Suite

Development Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_800d68d6/skill-function-test.

About this skill

Problem

Skill updates can fail after edits: broken references, boundary crashes, skipped configurations, and fixes without regression evidence. skill-function-test turns these risks into checkable stages: it backs up the target skill, builds a blueprint and constraint list from SKILL.md and scripts/, then runs scenario tests, functional tests, and S4 execution-fidelity tests before writing test-report.md and an HTML report. It treats configuration as the workflow, reducing prompts to ask users whether to enable fixes or skip steps.

Workflow and limits

The core is a dual-track testing model:
- S1-S3: checks trigger phrases, core capabilities, and multi-step workflows against declared behavior.
- D1-D6: inspects AST syntax, reference chains, data contamination, noise, calculation correctness, and boundary robustness.
- S4: validates rule fidelity under L1-L5 noise, including challenges, skip requests, reverse instructions, environmental contamination, and condition changes.

The flow is orchestrated by runner.py, with backup.py for ZIP backups, test_config.py for configuration, hooks.py for blocking, and gen_report.py for dual-format reports. Regression confirmation is required before accepting fixes, so F-0 BLOCK issues are not silently masked. Scope: it targets structured skill packages, not plain code snippets, manual tests, or ordinary unit tests; if the target lacks SKILL.md, a scripts directory, or parseable entry points, the scan and conclusions may be less reliable.

Use Cases

  • Before updating a multi-script skill, back it up and scan its blueprint to verify entry points, references, and constraints.
  • After an F-0 blocker appears, use S1-S3 and D1-D6 to locate failures in triggers, workflows, and functions.
  • Check whether a skill follows its rules under user challenges or skip requests, producing S4 trace records.
  • After fixes, run regression and export Markdown/HTML reports with conclusions written to the target skill's test-report.md.

Best For

  • Skill maintainers who edit script-heavy skills and want backups, scans, and traceable tests before release.
  • Automation engineers auditing LLM workflows for config and rule adherence, with focus on hooks and S4 fidelity.
  • Engineering leads inheriting skill packages who need to debug broken references, boundary crashes, and missing reports.
  • QA engineers adding skill tests to release checks and requiring F-0 fixes plus Markdown/HTML evidence.