Introduction

The core philosophy of DeepSeek Harness is “everything is a plugin.” Developers can extend an agent’s data processing capabilities through plugins. In local tariff data analysis scenarios, common pain points include inconsistent raw data structures, difficult-to-reuse analysis scripts, and hard-to-trace quality validation. To address these needs, this article introduces the dsh-plugin-csv-report plugin, which encapsulates the data preparation and descriptive statistics workflows as two dsh tools.

What Is It

The plugin is maintained by SUFE-Chaoyi and aims to refactor the tariff data analysis workflow into an installable, reusable, and auditable DSH plugin. It addresses inconsistent raw table structures, difficulty reusing analysis scripts, hard-to-trace quality issues, and missing upstream/downstream hash chains for results.

Core Features

The plugin provides the following core capabilities:

  • Data Preparation (tariff_prepare): Supports reading XLSX, XLSM, and CSV files within dataDir; performs standardization (user identifier, activation date, numeric fields); generates a field dictionary with 19 fields; validates required fields, duplicate users, invalid numeric values, and brand enumeration values; outputs standardized CSV and a preparation-stage manifest.
  • Descriptive Statistics (tariff_describe): Generates field-by-field data quality reports; generates statistics for numeric fields such as mean, standard deviation, median, and quartiles; generates Top N distributions for categorical fields; generates binary feature summaries for value contracts, external-network primary cards, and broadband; outputs grouped statistics by package name, user location, and brand series; generates a report manifest and Markdown index.
  • Data Security: The tools only read relative paths within dataDir and return no user-level records.
  • Reproducibility: Uses SHA-256 to link raw files, standardized data, and reports, and locks Node.js and Python dependencies.

Installation and Enablement

  1. Ensure dsh, Node.js, pnpm, and Python 3 are installed.
  2. Run the installation command:
    dsh plugin --profile demo add github:SUFE-Chaoyi/dsh-plugin-csv-report
  1. When installing from GitHub, pnpm may require you to allow the plugin to run the prepare build script. Only if you trust the source code, add the following to the profile’s pnpm-workspace.yaml:
    allowBuilds:
      dsh-plugin-csv-report: true
  1. Run the installation again and verify:
    dsh --profile demo --dump-config

Typical Usage

After installation, configure the paths in cordis.patch.yml for your dsh profile. The core configuration items are as follows:

- insert:
    - id: csv-report
      name: dsh-plugin-csv-report
      config:
        dataDir: /absolute/path/to/data
        outputDir: /absolute/path/to/reports
        pythonPath: /absolute/path/to/python
        defaultSheet: 数据
        dictionarySheet: 字段说明
        preparedSubdir: prepared
        minimumGroupSize: 30
        topCategoryLimit: 20
        maxInputBytes: 52428800
        timeoutMs: 120000

After configuration, the typical usage flow is as follows:

  1. Data Preparation: Place the raw tariff table in the dataDir/raw/ directory and invoke tariff_prepare in the dsh conversation.
    请调用 tariff_prepare,处理 raw/老旧资费特征.xlsx。
    主数据工作表是“数据”,字段说明工作表是“字段说明”。
    请返回数据质量状态、样本量、字段数、标准化 CSV 相对位置和警告摘要。
After successful processing, the standardized file is located at `dataDir/prepared/<原文件名>-<原始文件哈希前12位>/normalized_tariff.csv`.
  1. Descriptive Statistics: Use the relative path generated in the previous step and invoke tariff_describe.
    请调用 tariff_describe,分析:
    prepared/<原文件名>-<哈希>/normalized_tariff.csv

    请只基于聚合结果总结样本规模、数据质量、套餐费用、话费、流量使用、分类构成和分组差异;不要输出用户级记录,不做运营决策。

tariff_describe will generate a set of files under outputDir, including data quality, numeric summaries, categorical distributions, binary feature summaries, grouped statistics, and a Markdown report.

Use Cases and Notes

  • Use Cases: Suitable for scenarios requiring local tariff data processing, standardized field definitions, generating descriptive statistical reports, and ensuring data reproducibility.
  • Notes:
    • This plugin does not automatically make business decisions, and does not provide causal inference or recommendations for user operations actions.
    • You must trust the source code to allow the pnpm prepare script to run.
    • The data security boundary is to only read relative paths within the specified dataDir.
    • Data identifiers are used only for deduplication and quality validation; no user-level records are returned.

After the steps above, the data preparation and descriptive statistics workflows are encapsulated as standardized tool calls. This plugin ensures the reproducibility of analysis results by locking environment dependencies and performing hash validation. For the related directory and source code, see: Plugin Directory | GitHub Repository.