AI Agent Hub
Back to skills
📊

Exploratory Data Analysis

Data Analysis Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @org-02qudk26/exploratory-data-analysis by following https://skillhub.cn/install/skillhub.md.

About this skill

Problem It Solves

Scientific data files are often opaque: after receiving .fastq, .h5ad, .czi, or .parquet, you still need to know the format, expected fields, Python libraries, data quality, and appropriate downstream analysis. Manually reading documentation, writing loaders, and computing summaries can miss format-specific details or produce a non-reproducible plan. exploratory-data-analysis turns this pre-analysis step into a structured EDA workflow.

How It Works

The skill takes a file path and focuses on format detection, metadata extraction, quality checks, and report generation:
- File type detection: maps extensions to chemistry/molecular, bioinformatics, imaging, spectroscopy, omics, or general scientific data.
- Format lookup: reads the matching references/*.md file to get typical data, use cases, Python libraries, and EDA methods, such as pandas, h5py, or biopython.
- Analysis execution: for tables it checks dimensions, missing values, duplicates, and outliers; for sequences it inspects lengths, GC content, and quality scores; for images it checks X/Y/Z/C/T, bit depth, and metadata; for arrays it checks shape, dtype, and invalid values.
- Report output: produces a Markdown report using assets/report_template.md, covering metadata, format description, statistical summary, key findings, preprocessing advice, visualization, and next analysis steps.

Boundaries and Caveats

It is best for pre-analysis triage and planning, not full modeling or deep statistical testing. Large files should be sampled, chunked, or memory-mapped. Unknown extensions may need user clarification. Domain-specific libraries such as pysam, rdkit, or tifffile may need to be installed first. For multi-file work, run EDA per file and then compare, rather than assuming one format-specific pipeline fits all inputs.

Use Cases

  • Check FASTQ length, GC, quality scores.
  • Inspect .h5ad shape and expression stats.
  • Read .czi dims, channels, metadata.
  • Audit CSV missing, duplicate, outlier.

Best For

  • Bioinformatics engineer: FASTQ format and quality checks.
  • Research chemist: .pdb/.cif fields and EDA methods.
  • Data engineer: .parquet schema, distributions, gaps.
  • Graduate scientist: .czi dimensions and intensity stats.