AI Agent Hub
Back to skills
Data Governance Autopilot icon

Data Governance Autopilot

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md and install @user_69009747/data-governance-autopilot.

About this skill

Problem

Data governance often stalls because metadata is scattered, lineage is missing, and quality rules are maintained manually. Documentation drifts after schema changes, sensitive fields rely on heuristic labels, and audit evidence is assembled late. This skill turns governance into an executable pipeline with YAML, collectors, checkers, and audit artifacts instead of documentation alone.

How It Works

  • Metadata and lineage: multi-source collectors scan schemas, comments, and storage statistics; --incremental collects only changed tables. Automatic field/table lineage inference covers integration paths.
  • Quality checks: six-dimensional rules define constraints such as severity: critical; the engine runs checks and produces scores, helping maintain rules by priority.
  • Sensitive data: an AI classification engine grades sensitive fields. L3 partial masking preserves analytical value, while L4/L5 can use Tokenization instead of Masking.
  • Audit output: generates compliance reports and maps DAMA domains, showing coverage for data architecture, security, metadata, and quality.

Boundaries

The skill focuses on architecture, modeling, security, quality, and metadata automation. Storage, integration, and master data are only partially covered. Classification accuracy depends on --threshold; mislabels may need manual feedback. Masking choices should balance downstream analytics.

Use Cases

  • Collect changed metadata and infer field lineage after schema changes to trace downstream reporting impact.
  • Run six-dimensional DW quality checks before release, filter critical fields, and generate scored evidence.
  • Grade L3-L5 sensitive fields before audits, apply masking or Tokenization, and export a compliance report.
  • Slow metadata scans on thousands of tables can use --incremental to scan only last_modified changes.

Best For

  • Data engineers maintaining metadata, lineage, and storage statistics after multi-source schema changes.
  • Quality analysts using YAML rules, severity levels, and quality scores to triage critical issues.
  • Security and compliance owners grading sensitive fields, choosing masking policies, and preparing audit evidence.
  • DW leads checking referential integrity, quality rules, and governance coverage across warehouse layers.