Shopping Agent Experience Evaluation
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_d9b3e84b/shopping-agent-evaluation into your AI assistant.
About this skill
Problem
Shopping assistants can answer product recommendations, budget, brand preference, and feature questions, but single-turn impressions are not enough to compare Doubao and Qwen fairly. This skill provides a repeatable, quantified evaluation workflow for real e-commerce scenarios.
How It Works
- Query generation: Uses
DeepSeekto create 30 user queries covering 11 shopping scenarios, including specific products, budget constraints, brand preferences, functional requirements, ambiguous requests, and conflicting needs. - Model responses: Runs
doubao-seed-2-0-proandqwen-plusas shopping assistants to answer all queries. - Expert scoring: Uses
DeepSeekto score each answer on a 7-point scale across 11 dimensions, such as product relevance, persuasive recommendations, price reasonableness, factual accuracy, and conversational naturalness. - Deliverables: Produces a
DOCXreport with donut and radar charts, plus anXLSXfile containing queries, responses, dimension scores, and summary statistics.
Boundaries and Notes
It is best used for comparing shopping assistant experience, not for production-scale system load testing. The workflow makes many API calls and may run for 15–30 minutes, so DEEPSEEK_API_KEY, DOUBAO_API_KEY, and QWEN_API_KEY must be configured. Scores reflect the fixed query set and scoring prompts; validate with business-specific data before using them for model selection.
Use Cases
- Before selecting Doubao or Qwen for an e-commerce assistant, compare their answers across 30 shopping queries and generate 11-dimension scores.
- Turn shopping-agent test results into a DOCX report with radar charts, donut charts, and model pros and cons.
- Audit AI shopping guidance for budget, brand preference, and ambiguous requests to find off-topic, unsafe, or weakly explained answers.
- Export model responses, 11-dimension scores, success rates, and error causes into XLSX for review.
Best For
- Product managers who need to compare Doubao and Qwen answers for AI shopping assistant acceptance
- Evaluation engineers who need to package e-commerce assistant benchmarks into a DOCX report for leadership
- Prompt engineers who need to locate weak recommendation logic and factual errors in shopping assistant outputs
- Technical leads who need archived score details and error statistics for AI vendor comparison
Related Skills
Automatically searches job postings based on the user profile, AI-scores fit, saves desktop reports, and sends email updates with scheduled tracking.
Activates bionic reasoning for causal judgment, numerical prediction, and hypothesis validation, using hypothesis-driven checks, Bayesian updates, falsifiability tests, bias defense, and physical constraints.
Triggered by /plan, it asks the agent to output a plan, risks, impact scope, and validation approach before execution.
Track and clean agent session files, packages, and Skills via trash-first safety.