Merchant Data Ingestion
Paste the following prompt into your AI chat to install this skill:
Please install @user_223dc0b0/ksmy003 following https://skillhub.cn/install/skillhub.md.
About this skill
Problem
Collected web data often sits in ad hoc tasks, logs, or intermediate files, while raw html and parsed json lack a consistent persistence target. This makes it hard to identify when a page structure changed, and increases the risk of overwritten records or inconsistent fields. When page structure, API payloads, or merchant data change frequently, temporary files are not a reliable long-term reference.
How It Works
The skill turns ingestion into a concrete write pipeline:
- Database connection: connects to PostgreSQL using provided credentials.
- Data preparation: assembles merchant_id, url, html, json, and a timestamp.
- Targeted writes: stores raw HTML in merchant_raw_pages, and structured JSON in business tables or JSONB fields.
- Versioned snapshots: checks whether a snapshot exists for the day, then creates a new version record when needed for traceability.
- Result feedback: returns success, failure, and error details for downstream task handling.
Boundaries
It does not replace cleaning rules; it persists the cleaned raw page and structured result by merchant dimension. For long HTML or high-volume writes, account for PostgreSQL storage limits, index pressure, and write performance. Sensitive fields should be masked before insertion.
Use Cases
- After merchant pages are collected, write raw HTML and JSON to PostgreSQL by merchant_id.
- After page parsing, store order JSON in business tables or JSONB fields for later queries.
- When tracing same-day merchant data changes, create a snapshot and keep queryable version records.
- When a write fails, inspect the returned error to locate connection, credential, or storage-limit issues.
Best For
- Engineers handling merchant data ingestion, who need stable PostgreSQL writes for HTML and JSON.
- Analysts doing data traceability, who need to query merchant snapshots by date.
- Developers maintaining ETL pipelines, who need to confirm structured results land in business tables.
- Platform engineers handling sensitive web data, who need masking and storage-pressure control before insertion.
Related Skills
Analyzes smart customer-service chat logs with semantic clustering to generate word clouds, Top K frequent questions, and personalized guess-you-ask recommendations.
Maps natural-language Reddit requests to KeyAPI REST workflows, validates endpoint contracts against official docs, and executes search, detail, comment, ranking, and report tasks.
Schedules fetching from NEP and custom academic sites, filters and ranks papers by keywords, generates Chinese summaries, and pushes Feishu cards with local download archiving.
Browser-based CSV/TXT visualizer with multi-Y-axis curves, smoothing, and PNG/JSON/CSV export.