Data Cleaning & Format Conversion Tool
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_176cb31c/data-clean-transform.
About this skill
Problem It Solves
Data often becomes difficult to use before analysis or modeling begins: CSV, JSON, XML, YAML, Excel, and TSV files are mixed together, headers are inconsistent, and dirty rows contain nulls, duplicates, malformed values, encoding issues, and inconsistent formats. Manual cleanup is easy to repeat wrong, while ad hoc scripts usually fit only one file shape. This skill turns common messy-data problems into a reusable set of cleaning and conversion operations.
How It Works
It is organized around a file-in, rule-processing, output workflow. Key capabilities include:
- Format conversion: move structured data between
CSV,JSON,XML,YAML,Excel, andTSV. - Data cleaning: handle duplicates, null values, abnormal-value repair, and format normalization.
- Encoding repair: detect garbled text and convert between encodings such as
UTF-8,GBK,GB2312, andLatin1. - Regex processing: extract, replace, or split column values using regular expressions.
- Column operations: rename, map, split, or merge columns and apply type conversions.
- Data validation: check email addresses, phone numbers, ID numbers, addresses, and other field formats.
- Batch processing: convert and clean files at the directory level, useful for many files with similar structures.
Boundaries and Notes
This skill fits local file or directory data cleanup, format unification, and field standardization, especially when messy files need to become usable tabular data before analysis. It is not a real-time pipeline, a full ETL orchestrator, or a database governance system. If the task includes streaming data, transactional database operations, cross-system synchronization, or access control, use this as one processing step in a larger pipeline. For sensitive fields, define masking and retention rules before running batch cleanup.
Use Cases
- Before analysis, unify vendor CSV, Excel, and JSON files into one tabular structure.
- Clean duplicate rows, null values, and invalid phone numbers in a feedback sheet for import.
- Repair garbled text in GBK or Latin1 files and batch convert them to UTF-8 CSV.
- Use regex to split an address column into province, city, and district, then validate emails and IDs.
Best For
- Data analysts: preparing multi-source files into consistent CSV or Excel tables before modeling.
- Backend engineers: fixing log or API export encoding issues, splitting fields, and validating formats.
- Operations analysts: consolidating vendor or campaign data by removing duplicates, handling nulls, and converting files.
- ETL script maintainers: turning file-cleaning rules for fixed structures into reusable processing steps.
Related Skills
Generate Markdown public opinion reports by calling an internal service with MIDU_API_KEY.
Extracts Google AI Mode answers, standard SERP, AI Overviews, and citations via Pangolin APIs, with multi-turn follow-ups and region support.
Provides break-even analysis frameworks and templates without code execution, outputting structured recommendations.
Generate web reports from existing analysis data with classic or PPT-style layouts, Chart.js charts, and keyboard/touch navigation.