AI Agent Hub
Back to skills
Big Data Management and Applications icon

Big Data Management and Applications

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Install @user_a045eab5/bigdata-management according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem it solves

When a request mentions real-time dashboards, data-lake migration, or Spark tuning but lacks data volume, latency, and team stack, the answer can collapse into a tool list. This skill first constrains the inputs: business context, focus area, stack preference, and whether the question is theoretical, implementation, or design.

How it works and where it stops

  • Collection and storage: It compares Kafka, Flume, DataX, HDFS, Ceph, Iceberg, Hudi, and Delta Lake for ingestion and persistence choices.
  • Processing and query: It maps Spark, Flink, Kafka Streams, Airflow, ClickHouse, Doris, Trino, and Druid to batch, stream, scheduling, and OLAP concerns.
  • Governance and security: It brings in Atlas, DataHub, Great Expectations, and Ranger for catalog, lineage, quality, and access control.

It expects actionable output: trade-offs, tool comparisons, key parameters, text architecture diagrams, and code examples when needed. Boundaries: it is closer to engineering consulting and solution design than production operations, cost accounting, or benchmarking. For emerging components or private implementations, verify against official docs and community cases.

Use Cases

  • Design a 1B-event daily real-time pipeline and define the roles of Kafka, Flink, and the OLAP engine.
  • Migrate a legacy Hive warehouse to Iceberg or Hudi with a phased migration and rollback plan.
  • Troubleshoot slow Spark SQL jobs, heavy shuffle, and data skew, then produce tuning checks for partitions, memory, and parameters.
  • Build quality validation and lineage tracking for risk-control analytics, covering schemas, metric definitions, and access control.

Best For

  • A backend engineer choosing a real-time warehouse stack who needs to compare Kafka, Flink, ClickHouse, or Doris.
  • A data engineer handling platform governance who needs catalog, lineage, quality rules, and permission policies.
  • An architect leading warehouse-to-lakehouse migration who needs to compare Iceberg, Hudi, and Delta Lake and plan the rollout.
  • A Spark developer optimizing job performance who needs to diagnose shuffle, partitioning, memory, and skew issues.