Skill Evaluation Pro
Paste the following prompt into your AI chat to install this skill:
Please install @user_616d1d5a/skill-evaluation-pro according to the guide at https://skillhub.cn/install/skillhub.md.
About this skill
Problem
When evaluating local Skill behavior, the hard part is not whether a run succeeds, but how to compare candidates consistently: different models, frameworks, and multiple Skill variants can produce unstable results if evaluation sets, drivers, and judging criteria are assembled by hand. Skill Evaluation Pro turns a natural-language evaluation intent into a controlled workflow, supporting single-Skill validation, multi-Skill comparison, multi-model comparison, and runtime-framework comparison.
How it works
The process advances conversationally. First, it confirms the evaluation scenario; next, it selects the target Skill, driver model, and judge model; then it uploads an existing evaluation set or generates one automatically; finally, it presents a task summary and scoring report. Key steps include:
- Scenario confirmation: asks only for the scenario before introducing model or dataset choices.
- Configuration: selects the
Skill, driver model, and judge model, then confirms the setup. - Data preparation: uploads a set or generates data, then previews it for confirmation.
- Execution and reporting: handles auth, packaging, upload, and polling silently, then surfaces the result summary.
Boundaries
This skill is for automated effect evaluation of local Skills. It does not replace code review, system performance stress testing, or model training/fine-tuning. Interaction stays in Chinese, uses numbered tables and concise summaries, and avoids exposing internal commands, fields, or IDs. If the goal includes stress metrics or training pipelines, use a separate workflow.
Use Cases
- Before choosing among local Skills, compare their output quality on the same task.
- When integrating one Skill with different models, check performance differences across driver models.
- Compare multiple runtime frameworks on the same evaluation set to verify stable Skill behavior.
- Quickly validate one local Skill against expected outputs before release.
Best For
- AI engineers responsible for local Skill quality acceptance who need consistent model-output comparison.
- Developers maintaining multi-framework agent apps who need to verify the same Skill across runtimes.
- Product or algorithm engineers comparing candidate Skills who need task-based side-by-side scoring.
- LLM application engineers preparing model switches who need to validate a Skill on target models.
Related Skills
Model routing, persistent parameter management, self-check repair, and global default model control for XiaoYi Claw.
Install, update, and manage OpenClaw Skills through SkillHub, with automatic detection after installation.
Provides API endpoints for AI agents to post bottles and graffiti, browse the feed, and interact with likes and comments.
ReqPlan constrains agent development, debugging, and analysis workflows with a seven-stage state machine, checkpoints, quality audits, and local Harness artifacts.