AI Agent Hub
Back to skills
SkillsBench Skill Evaluator icon

SkillsBench Skill Evaluator

Development Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_f37e97ba/skillsbench-evaluator.

About this skill

Problem Solved

Many agent skill repositories have accumulated large numbers of SKILL.md files, but there is no repeatable way to judge quality: which skills trigger reliably, which merely read smoother, which may mislead the model, and which reduce errors. Manual review is hard to quantify, and large skill libraries need a common evaluation baseline.

skillsbench-evaluator treats skills as comparable assets using a fixed scoring rule. It does not rely on subjective impressions; each dimension must cite SKILL.md evidence before a score, a tier, and a fix are reported.

How It Works

  • Single-skill review: identify the skill directory, read the full SKILL.md and its references/ structure, then score the 12 dimensions in references/rubric.md as 0 or 1.
  • Evidence-based scoring: annotate each dimension with the supporting original text, avoid guessing intent, and map the total score to Foundation, Professional, or Premium tiers.
  • Batch review: scan ~/.workbuddy/skills/, then output a ranked table, a dimension heatmap, and recurring cross-skill defects.
  • Actionable fixes: generate repair suggestions for failed dimensions, prioritizing high-risk items such as trigger conditions, instruction structure, gotchas, and tool controls.

It fits skill-library governance, pre-release review, and comparing similar agent skills. Keep in mind that results depend heavily on the match between SKILL.md and rubric.md, so rule or template changes can shift scores. It is better for judging whether skill documentation is executable than for replacing real A/B testing or production validation.

Use Cases

  • Before publishing a skill library, review each SKILL.md in one pass for trigger conditions, instruction structure, tool controls, and report concrete gaps.
  • Maintain multiple agent skills for weekly governance by scanning all SKILL.md files, generating a ranked table, dimension heatmap, and recurring defects.
  • Review a workflow skill for release readiness by scoring 12 rubric dimensions with original-text evidence and producing actionable repair suggestions.
  • Compare two tool-type skills before routing work by checking trigger conditions, gotchas, and constraints to identify the more stable invocation path.

Best For

  • Engineering owners publishing an agent skill library who need a release gate to check executable SKILL.md files before cross-team rollout.
  • Developers maintaining workflow skills who need to compare trigger conditions, instruction structure, and gotcha descriptions across versions before changing prompts.
  • Platform leads governing an enterprise skill directory who need to fold scattered SKILL.md files into a unified scoring table and rank gaps.
  • Engineers writing tool-type skills who need to identify which failed dimensions affect invocation success, error rate, or safety before shipping.