AI Agent Hub
Back to skills
Lakehouse Architect icon

Lakehouse Architect

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_69009747/lakehouse-architect.

About this skill

Problem

When a company maintains both a warehouse and a data lake, it often faces duplicate storage, inconsistent metrics, and fragmented query paths. This skill breaks lakehouse design into practical steps, helping teams choose among Iceberg, Hudi, Paimon, and Delta Lake, then build an end-to-end architecture around Flink, Spark, and StarRocks.

How It Works

  • Format selection: compares real-time capability, Flink / Spark integration, upsert performance, and community maturity to narrow down scenarios.
  • Architecture design: explains lakehouse patterns and recommends Paimon + StarRocks for streaming ingestion and accelerated querying.
  • Implementation and migration: covers environment setup, Flink CDC ingestion, layered processing, external catalog queries, snapshot governance, dual-run, canary, and full cutover with rollback controls.
  • Cost analysis: uses a TCO calculator to compare storage, compute, operations, and redundancy costs before migration.

Boundaries

It fits data teams that already use MySQL, Flink, Spark, or StarRocks and need architecture evaluation or migration planning. If the org is still proving concepts or mainly needs business reporting, basic data models and metric definitions should be established first.

Use Cases

  • Evaluate a dual warehouse-plus-lake setup and choose among Iceberg, Hudi, Paimon, and Delta Lake.
  • Design an ODS-DWD-ADS pipeline with Flink CDC and Paimon, then accelerate queries via StarRocks.
  • Plan warehouse-to-lakehouse migration with dual-run, canary cutover, consistency checks, and rollback.
  • Compare storage, compute, ops labor, and redundancy costs before and after migration for TCO review.

Best For

  • Data platform engineers responsible for data architecture who need to select a lake format and design a unified pipeline.
  • DBAs or data engineers leading warehouse migration who need dual-run, canary cutover, and rollback plans.
  • Operations leaders tracking compute and storage costs who need TCO comparison before migration.
  • Flink and StarRocks developers who need to ingest MySQL CDC data into a lakehouse and accelerate queries.