AI Agent Hub
Back to skills
Customer Insight Analyzer icon

Customer Insight Analyzer

Data Analysis Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_fc05d2ec/customer-insight-analyzer into your AI assistant.

About this skill

Problem It Solves

Smart customer-service chats are messy: users may express the same issue with different wording, colloquial phrasing, or typos. Product and operations teams usually need to know which questions are genuinely frequent, which segments care about what, and how to build a “guess you ask” recommendation list that is precomputed by user tags instead of relying on exact string matching. customer-insight-analyzer is aimed at this kind of chat-log analysis, turning unstructured service conversations into sortable, renderable, and segment-aware recommendation data.

How It Works

The skill accepts chat records in CSV or JSON format and relies primarily on semantic vectors rather than keyword exact matching. Its main steps are:
- Data parsing: read the input file and normalize records for analysis;
- Chinese tokenization: use jieba for word-level splitting, supporting word-cloud generation and stopword handling;
- Semantic embedding: default to the all-MiniLM-L6-v2 model, documented as roughly 80MB and runnable on CPU locally;
- Clustering: use HDBSCAN density clustering and let the algorithm determine the number of clusters;
- Result generation: output word-cloud data, frequent-question Top K lists, and precomputed “guess you ask” lists grouped by user tags.

From an engineering perspective, it is better suited to offline batch analysis: process a set of chat records and emit structured JSON output that a frontend can use to render word clouds or recommendation widgets. If the default Chinese semantic model performs poorly, the documentation suggests replacing the model in config.yaml; the exact replacement should depend on the models available in your runtime environment.

Boundaries and Caveats

  • It assumes the input is customer-service chat data and that the fields match the documented input format;
  • First run may download a model, so the environment needs access to Hugging Face or ModelScope;
  • It is a local CPU pipeline without GPU dependency, but processing time grows with data volume; the referenced benchmark is about 10 seconds for 3,500 records, 60 seconds for 25,000 records, and 5 minutes for 100,000 records;
  • Personalized “guess you ask” recommendations depend on user tags in the input, so missing or coarse tags can reduce segmentation quality.

Use Cases

  • CS ops converts a week of CSV chats into word clouds and Top K frequent questions to spot repeated entry points.
  • Product precomputes tag-grouped guess-you-ask lists from JSON sessions for frontend recommendation widgets.
  • ML engineers inspect HDBSCAN clusters to diagnose fragmented user phrasing and evaluate swapping the embedding model.
  • Data analysts run CSV service logs as a local CPU job and export analysis_result.json for offline review.

Best For

  • CS ops: needs high-frequency chat topics, word clouds, and Top K lists for service reviews.
  • Product managers: needs tag-based guess-you-ask recommendations for frontend widgets.
  • Data analysts: needs offline semantic clustering of CSV/JSON service logs and JSON exports.
  • Frontend/backend engineers: needs structured word-cloud and recommendation data for service UI modules.