Multi-Source Data Cleaner Pro
Paste the following prompt into your AI chat to install this skill:
Please install @user_e9af5021/multi-source-data-cleaner-pro according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem it solves
When teams deliver CSV, Excel, JSON, Parquet, or SQL dumps that must be combined into one analytical table, the bottleneck is often not reading files but reconciling inconsistent column names, mixed encodings, varied date and boolean formats, different missing-value semantics, ambiguous duplicates, and inconsistent PII handling. multi-source-data-cleaner-pro targets this batch preprocessing task and turns ad hoc cleaning rules into an auditable pipeline.
How it works
- Source profiling: detects format, encoding, delimiter, header row, and infers column types, null rates, and cardinality.
- Type normalization: converts thousands separators, scientific notation, and currency symbols to numeric values; normalizes many date formats to ISO 8601; standardizes booleans, phone numbers to
E.164, Chinese full/half-width names, and ID prefixes. - Missing-value handling: applies column-level
drop_row, mean/median/mode imputation, constant fills,forward_fill, or interpolation, and writes an_imputedflag. - Schema reconciliation: aligns columns by exact match, fuzzy match, and type compatibility, producing
crosswalk.json; users can override withmapping.yaml. - Fuzzy deduplication: uses blocking and record linkage for names, addresses, and IDs, with lower-confidence merges routed for human review.
- Quality scoring: rates completeness, accuracy, consistency, timeliness, uniqueness, and validity from 0 to 100 with drill-down details.
Boundaries
It is intended for offline batch work on structured or semi-structured tabular data, not free-text extraction, proprietary binary formats without parsers, or real-time stream cleaning. PII is masked by default, original files are left untouched, and processing stays local.
Use Cases
- Unify CSV, Excel, and JSON exports from multiple teams into one customer table
- Clean an order table with mixed dates, booleans, and phone formats, flagging imputed values
- Deduplicate customer records with fuzzy matching and keep low-confidence merges for review
- Produce a six-dimension data-quality scorecard with completeness, consistency, and uniqueness details
Best For
- Data analyst consolidating exports from multiple systems: align column names and produce an auditable mapping
- Product operations managing customer master data: deduplicate names, phones, and addresses with PII masked
- Backend engineer building ETL: batch-normalize CSV, Excel, JSON, and Parquet into unified types
- Risk-control auditor checking data quality: score six dimensions and locate missing, inconsistent, duplicate records
Related Skills
Generate Markdown public opinion reports by calling an internal service with MIDU_API_KEY.
Extracts Google AI Mode answers, standard SERP, AI Overviews, and citations via Pangolin APIs, with multi-turn follow-ups and region support.
Provides break-even analysis frameworks and templates without code execution, outputting structured recommendations.
Generate web reports from existing analysis data with classic or PPT-style layouts, Chart.js charts, and keyboard/touch navigation.