Exam Quality Evaluator
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_638acef9/exam-evaluator.
About this skill
Problem it addresses
Exam evaluation often mixes two different jobs: objective metrics such as item count, answer distribution, difficulty, and duplicate detection need stable computation, while subjective quality—stem clarity, distractor design, explanation quality, and cognitive coverage—needs semantic judgment. A simple weighted average can hide weak dimensions, while manual review is hard to reproduce. exam-evaluator focuses on exam-paper and question-bank quality assessment, converting Excel, Word, and PDF materials into structured data and producing an HTML report.
How it works and where to be careful
The workflow runs in phases. Phase 1 extracts text locally: Excel paths can emit clean.json directly, while Word and PDF paths produce raw.json for agent structuring. The agent identifies item type, stem, options, answer, explanation, knowledge point, cognitive level, and difficulty. Phase 2+ validates answer coverage, item-type coverage, and knowledge-point coverage, and flags low-quality parses for secondary review. Phase 3 runs compute_metrics.py and detect_duplicates.py for quantitative metrics. Phase 4 requires user confirmation of evaluation mode, weight preset, and low-confidence items. Phase 5 scores five dimensions—content validity, structural validity, difficulty control, discrimination potential, and presentation norms—using the weak-point rule min(weighted average, lowest score + 2). Phase 6 renders an HTML report with Chart.js charts, item-type distribution, difficulty distribution, answer distribution, and knowledge-point visualization.
Use with caution: it does not replace curriculum review and may be less reliable on encrypted files, poor scans, or noisy OCR text. Fill-in-the-blank uniqueness, low-confidence item types, and skewed answer distributions should be confirmed or annotated by the user.
Use Cases
- Review a final-exam Word file for item count, answer distribution, difficulty gradient, duplicates, and an archivable HTML report.
- Check a question-bank Excel file for knowledge coverage, true/false ratio, skewed multiple-choice answers, and low-confidence items.
- Split multiple independent exams into separate clean JSON files, compute metrics for each, and build an index page.
- Evaluate a placement mock exam using placement weights for difficulty, discrimination, and cognitive level, then output prioritized fixes.
Best For
- Curriculum leads who need to audit exam structure, answer distribution, and presentation standards.
- Exam authors who convert Word/PDF papers into structured banks and quality reports.
- Academic data analysts responsible for knowledge coverage and duplicate detection.
- Subject leads who need placement, routine, or competition-weighted evaluation of mock exams.
Related Skills
Generate Markdown public opinion reports by calling an internal service with MIDU_API_KEY.
Extracts Google AI Mode answers, standard SERP, AI Overviews, and citations via Pangolin APIs, with multi-turn follow-ups and region support.
Provides break-even analysis frameworks and templates without code execution, outputting structured recommendations.
Generate web reports from existing analysis data with classic or PPT-style layouts, Chart.js charts, and keyboard/touch navigation.