Data Governance Autopilot
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md and install @user_69009747/data-governance-autopilot.
About this skill
Problem
Data governance often stalls because metadata is scattered, lineage is missing, and quality rules are maintained manually. Documentation drifts after schema changes, sensitive fields rely on heuristic labels, and audit evidence is assembled late. This skill turns governance into an executable pipeline with YAML, collectors, checkers, and audit artifacts instead of documentation alone.
How It Works
- Metadata and lineage: multi-source collectors scan schemas, comments, and storage statistics;
--incrementalcollects only changed tables. Automatic field/table lineage inference covers integration paths. - Quality checks: six-dimensional rules define constraints such as
severity: critical; the engine runs checks and produces scores, helping maintain rules by priority. - Sensitive data: an AI classification engine grades sensitive fields.
L3partial masking preserves analytical value, whileL4/L5can use Tokenization instead of Masking. - Audit output: generates compliance reports and maps DAMA domains, showing coverage for data architecture, security, metadata, and quality.
Boundaries
The skill focuses on architecture, modeling, security, quality, and metadata automation. Storage, integration, and master data are only partially covered. Classification accuracy depends on --threshold; mislabels may need manual feedback. Masking choices should balance downstream analytics.
Use Cases
- Collect changed metadata and infer field lineage after schema changes to trace downstream reporting impact.
- Run six-dimensional DW quality checks before release, filter critical fields, and generate scored evidence.
- Grade L3-L5 sensitive fields before audits, apply masking or Tokenization, and export a compliance report.
- Slow metadata scans on thousands of tables can use --incremental to scan only last_modified changes.
Best For
- Data engineers maintaining metadata, lineage, and storage statistics after multi-source schema changes.
- Quality analysts using YAML rules, severity levels, and quality scores to triage critical issues.
- Security and compliance owners grading sensitive fields, choosing masking policies, and preparing audit evidence.
- DW leads checking referential integrity, quality rules, and governance coverage across warehouse layers.
Related Skills
Analyzes smart customer-service chat logs with semantic clustering to generate word clouds, Top K frequent questions, and personalized guess-you-ask recommendations.
Maps natural-language Reddit requests to KeyAPI REST workflows, validates endpoint contracts against official docs, and executes search, detail, comment, ranking, and report tasks.
Schedules fetching from NEP and custom academic sites, filters and ranks papers by keywords, generates Chinese summaries, and pushes Feishu cards with local download archiving.
Browser-based CSV/TXT visualizer with multi-Y-axis curves, smoothing, and PNG/JSON/CSV export.