AI Agent Hub
Back to skills
ETL Pipeline Generator icon

ETL Pipeline Generator

Data Analysis Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md and install @user_15292d5a/yjkj-etl-pipeline-generator.

About this skill

Data Ingestion for Knowledge Graphs

Raw records often come from CSV, JSON, databases, REST APIs, and streaming sources with inconsistent fields, unclear entity boundaries, and inferred relationships. Loading that data directly into a graph store can corrupt the schema and degrade queries. The ETL Pipeline Generator breaks the workflow into configurable, rerunnable Extract, Transform, and Load stages, reducing ad-hoc scripting and troubleshooting.

How It Works

It generates a structured pipeline design from data sources, transformation rules, target systems, and quality requirements:
- Extract: connect files, databases, APIs, warehouses, or streaming sources to form a raw data stream.
- Transform: apply entity detection, relationship inference, schema mapping, deduplication, type conversion, validation, and filtering to create graph-ready structures.
- Load: write to Neo4j, RDF Triple Stores, ArangoDB, TigerGraph, or CSV/JSON targets.

Key artifacts include YAML pipeline configuration, DAG definitions, Python/Scala/SQL scripts, data flow diagrams, monitoring and alerting configs, and documentation. The execution path typically covers config validation, connector initialization, extraction, transformation, quality checks, loading, load verification, and metrics collection.

Boundaries and Notes

This skill is best for pipeline design and configuration generation, especially for automated ingestion, knowledge graph population, and data quality validation. It does not replace business semantics: entity definitions, relationship rules, and constraints must still be defined by the team. For high-throughput streaming, complex permission isolation, or production-grade disaster recovery, combine it with Airflow, Spark, Great Expectations, and similar systems, and add testing, version control, and credential management.

Use Cases

  • Extract customer data from CSV, JSON, and databases, then generate a rerunnable ETL config for Neo4j.
  • Split REST API order streams into entities and relationships, then output DAG and transformation rules for execution.
  • Generate schema mapping, deduplication, and validation steps for RDF triples before writing to a graph store.
  • Design monitoring, alerting, and error-recovery rules for multi-source ingestion, including row metrics and retries.

Best For

  • Data engineers owning knowledge graph ingestion: turn raw multi-source data into graph-ready structures and pipeline configs.
  • Backend engineers building graph applications: convert APIs, databases, and text into entities and relationships for Neo4j or RDF.
  • Data-quality analysts: add deduplication, type conversion, validation, and error handling to ETL to reduce dirty data.
  • Platform engineers automating data flows: produce DAGs, YAML, monitoring alerts, and docs for reuse and debugging.