AI Agent Hub
Back to skills
Statistical Data Format Converter icon

Statistical Data Format Converter

Data Analysis Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Install @user_ff7413f5/statdata-transfer by following https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Statistical and clinical-trial data are often locked across SPSS, Stata, SAS, Excel, Parquet, and similar files. A plain pandas read may preserve values but drop variable labels, value labels, special-missing-value meanings, and other metadata. Cross-format conversion makes this harder when tools disagree about which metadata can survive.

How It Works

The skill separates extraction from conversion. It first reads a source file into a pandas-compatible structure and reports which metadata is retained or lost. It then converts among common statistical formats or exports to Parquet, HDF5, JSON, CSV, Excel, and related targets. For binary statistical formats, it aims to preserve variable and value labels where supported. For Parquet and Arrow, labels can be embedded in schema.metadata; for CSV/TSV, a sidecar _metadata.json file helps keep metadata alongside the data. The default behavior is preview-and-warn: R-backed paths are opt-in and disabled unless allow_r_exec=True is explicitly set.

Boundaries

This is not a general-purpose ETL engine, and not every supported format is lossless. Some formats are detect-only, require optional R backends, or preserve only partial structure. Output should still be validated, especially before regulatory submission.

Use Cases

  • Ingest .sav or .dta survey files and convert them to Parquet while keeping value labels.
  • Turn SAS XPT output into Excel for review and identify which metadata fields may be lost.
  • Preview field labels and special-missing mappings before importing clinical CSV files into SPSS.
  • Audit label loss before CSV export and emit a sidecar _metadata.json file.

Best For

  • Data analysts merging SPSS or Stata results into Python workflows
  • Clinical data managers validating metadata before data delivery
  • Data platform engineers preserving statistical labels in ETL pipelines
  • Statistical researchers auditing format conversion for reproducible analyses