AI Agent Hub
Back to skills
Multi-Source Data Cleaner Pro icon

Multi-Source Data Cleaner Pro

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_e9af5021/multi-source-data-cleaner-pro according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem it solves

When teams deliver CSV, Excel, JSON, Parquet, or SQL dumps that must be combined into one analytical table, the bottleneck is often not reading files but reconciling inconsistent column names, mixed encodings, varied date and boolean formats, different missing-value semantics, ambiguous duplicates, and inconsistent PII handling. multi-source-data-cleaner-pro targets this batch preprocessing task and turns ad hoc cleaning rules into an auditable pipeline.

How it works

  • Source profiling: detects format, encoding, delimiter, header row, and infers column types, null rates, and cardinality.
  • Type normalization: converts thousands separators, scientific notation, and currency symbols to numeric values; normalizes many date formats to ISO 8601; standardizes booleans, phone numbers to E.164, Chinese full/half-width names, and ID prefixes.
  • Missing-value handling: applies column-level drop_row, mean/median/mode imputation, constant fills, forward_fill, or interpolation, and writes an _imputed flag.
  • Schema reconciliation: aligns columns by exact match, fuzzy match, and type compatibility, producing crosswalk.json; users can override with mapping.yaml.
  • Fuzzy deduplication: uses blocking and record linkage for names, addresses, and IDs, with lower-confidence merges routed for human review.
  • Quality scoring: rates completeness, accuracy, consistency, timeliness, uniqueness, and validity from 0 to 100 with drill-down details.

Boundaries

It is intended for offline batch work on structured or semi-structured tabular data, not free-text extraction, proprietary binary formats without parsers, or real-time stream cleaning. PII is masked by default, original files are left untouched, and processing stays local.

Use Cases

  • Unify CSV, Excel, and JSON exports from multiple teams into one customer table
  • Clean an order table with mixed dates, booleans, and phone formats, flagging imputed values
  • Deduplicate customer records with fuzzy matching and keep low-confidence merges for review
  • Produce a six-dimension data-quality scorecard with completeness, consistency, and uniqueness details

Best For

  • Data analyst consolidating exports from multiple systems: align column names and produce an auditable mapping
  • Product operations managing customer master data: deduplicate names, phones, and addresses with PII masked
  • Backend engineer building ETL: batch-normalize CSV, Excel, JSON, and Parquet into unified types
  • Risk-control auditor checking data quality: score six dimensions and locate missing, inconsistent, duplicate records