AI Agent Hub
Back to skills
Skill Evaluation Pro icon

Skill Evaluation Pro

AI Agent Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_616d1d5a/skill-evaluation-pro according to the guide at https://skillhub.cn/install/skillhub.md.

About this skill

Problem

When evaluating local Skill behavior, the hard part is not whether a run succeeds, but how to compare candidates consistently: different models, frameworks, and multiple Skill variants can produce unstable results if evaluation sets, drivers, and judging criteria are assembled by hand. Skill Evaluation Pro turns a natural-language evaluation intent into a controlled workflow, supporting single-Skill validation, multi-Skill comparison, multi-model comparison, and runtime-framework comparison.

How it works

The process advances conversationally. First, it confirms the evaluation scenario; next, it selects the target Skill, driver model, and judge model; then it uploads an existing evaluation set or generates one automatically; finally, it presents a task summary and scoring report. Key steps include:

  • Scenario confirmation: asks only for the scenario before introducing model or dataset choices.
  • Configuration: selects the Skill, driver model, and judge model, then confirms the setup.
  • Data preparation: uploads a set or generates data, then previews it for confirmation.
  • Execution and reporting: handles auth, packaging, upload, and polling silently, then surfaces the result summary.

Boundaries

This skill is for automated effect evaluation of local Skills. It does not replace code review, system performance stress testing, or model training/fine-tuning. Interaction stays in Chinese, uses numbered tables and concise summaries, and avoids exposing internal commands, fields, or IDs. If the goal includes stress metrics or training pipelines, use a separate workflow.

Use Cases

  • Before choosing among local Skills, compare their output quality on the same task.
  • When integrating one Skill with different models, check performance differences across driver models.
  • Compare multiple runtime frameworks on the same evaluation set to verify stable Skill behavior.
  • Quickly validate one local Skill against expected outputs before release.

Best For

  • AI engineers responsible for local Skill quality acceptance who need consistent model-output comparison.
  • Developers maintaining multi-framework agent apps who need to verify the same Skill across runtimes.
  • Product or algorithm engineers comparing candidate Skills who need task-based side-by-side scoring.
  • LLM application engineers preparing model switches who need to validate a Skill on target models.