Exploratory Data Analysis
Paste the following prompt into your AI chat to install this skill:
Please install @org-02qudk26/exploratory-data-analysis by following https://skillhub.cn/install/skillhub.md.
About this skill
Problem It Solves
Scientific data files are often opaque: after receiving .fastq, .h5ad, .czi, or .parquet, you still need to know the format, expected fields, Python libraries, data quality, and appropriate downstream analysis. Manually reading documentation, writing loaders, and computing summaries can miss format-specific details or produce a non-reproducible plan. exploratory-data-analysis turns this pre-analysis step into a structured EDA workflow.
How It Works
The skill takes a file path and focuses on format detection, metadata extraction, quality checks, and report generation:
- File type detection: maps extensions to chemistry/molecular, bioinformatics, imaging, spectroscopy, omics, or general scientific data.
- Format lookup: reads the matching references/*.md file to get typical data, use cases, Python libraries, and EDA methods, such as pandas, h5py, or biopython.
- Analysis execution: for tables it checks dimensions, missing values, duplicates, and outliers; for sequences it inspects lengths, GC content, and quality scores; for images it checks X/Y/Z/C/T, bit depth, and metadata; for arrays it checks shape, dtype, and invalid values.
- Report output: produces a Markdown report using assets/report_template.md, covering metadata, format description, statistical summary, key findings, preprocessing advice, visualization, and next analysis steps.
Boundaries and Caveats
It is best for pre-analysis triage and planning, not full modeling or deep statistical testing. Large files should be sampled, chunked, or memory-mapped. Unknown extensions may need user clarification. Domain-specific libraries such as pysam, rdkit, or tifffile may need to be installed first. For multi-file work, run EDA per file and then compare, rather than assuming one format-specific pipeline fits all inputs.
Use Cases
- Check FASTQ length, GC, quality scores.
- Inspect .h5ad shape and expression stats.
- Read .czi dims, channels, metadata.
- Audit CSV missing, duplicate, outlier.
Best For
- Bioinformatics engineer: FASTQ format and quality checks.
- Research chemist: .pdb/.cif fields and EDA methods.
- Data engineer: .parquet schema, distributions, gaps.
- Graduate scientist: .czi dimensions and intensity stats.
Related Skills
Fetches Baidu Hot Search Top 10 titles using web_fetch first, validates same-day data, and falls back to browser automation when stale.
Generates an evening A-share policy and trading opportunity daily report by collecting same-day index, policy, and capital data, then applying a fixed template to highlight beneficiary sectors, drivers, and price directions.
Maps natural-language TikTok requests to KeyAPI REST scenarios, validates endpoints against docs, and executes data queries and analysis.
Extract city-specified AI jobs from BOSS Zhipin, save CSV/table data, mark new postings, and summarize salary trends, application advice, and HTML reports.