Introduction¶
The core philosophy of DeepSeek Harness is “everything is a plugin.” Developers can extend an agent’s data processing capabilities through plugins. In local tariff data analysis scenarios, common pain points include inconsistent raw data structures, difficult-to-reuse analysis scripts, and hard-to-trace quality validation. To address these needs, this article introduces the dsh-plugin-csv-report plugin, which encapsulates the data preparation and descriptive statistics workflows as two dsh tools.
What Is It¶
The plugin is maintained by SUFE-Chaoyi and aims to refactor the tariff data analysis workflow into an installable, reusable, and auditable DSH plugin. It addresses inconsistent raw table structures, difficulty reusing analysis scripts, hard-to-trace quality issues, and missing upstream/downstream hash chains for results.
Core Features¶
The plugin provides the following core capabilities:
- Data Preparation (
tariff_prepare): Supports reading XLSX, XLSM, and CSV files withindataDir; performs standardization (user identifier, activation date, numeric fields); generates a field dictionary with 19 fields; validates required fields, duplicate users, invalid numeric values, and brand enumeration values; outputs standardized CSV and a preparation-stage manifest. - Descriptive Statistics (
tariff_describe): Generates field-by-field data quality reports; generates statistics for numeric fields such as mean, standard deviation, median, and quartiles; generates Top N distributions for categorical fields; generates binary feature summaries for value contracts, external-network primary cards, and broadband; outputs grouped statistics by package name, user location, and brand series; generates a report manifest and Markdown index. - Data Security: The tools only read relative paths within
dataDirand return no user-level records. - Reproducibility: Uses SHA-256 to link raw files, standardized data, and reports, and locks Node.js and Python dependencies.
Installation and Enablement¶
- Ensure dsh, Node.js, pnpm, and Python 3 are installed.
- Run the installation command:
dsh plugin --profile demo add github:SUFE-Chaoyi/dsh-plugin-csv-report
- When installing from GitHub, pnpm may require you to allow the plugin to run the
preparebuild script. Only if you trust the source code, add the following to the profile’spnpm-workspace.yaml:
allowBuilds:
dsh-plugin-csv-report: true
- Run the installation again and verify:
dsh --profile demo --dump-config
Typical Usage¶
After installation, configure the paths in cordis.patch.yml for your dsh profile. The core configuration items are as follows:
- insert:
- id: csv-report
name: dsh-plugin-csv-report
config:
dataDir: /absolute/path/to/data
outputDir: /absolute/path/to/reports
pythonPath: /absolute/path/to/python
defaultSheet: 数据
dictionarySheet: 字段说明
preparedSubdir: prepared
minimumGroupSize: 30
topCategoryLimit: 20
maxInputBytes: 52428800
timeoutMs: 120000
After configuration, the typical usage flow is as follows:
- Data Preparation: Place the raw tariff table in the
dataDir/raw/directory and invoketariff_preparein the dsh conversation.
请调用 tariff_prepare,处理 raw/老旧资费特征.xlsx。
主数据工作表是“数据”,字段说明工作表是“字段说明”。
请返回数据质量状态、样本量、字段数、标准化 CSV 相对位置和警告摘要。
After successful processing, the standardized file is located at `dataDir/prepared/<原文件名>-<原始文件哈希前12位>/normalized_tariff.csv`.
- Descriptive Statistics: Use the relative path generated in the previous step and invoke
tariff_describe.
请调用 tariff_describe,分析:
prepared/<原文件名>-<哈希>/normalized_tariff.csv
请只基于聚合结果总结样本规模、数据质量、套餐费用、话费、流量使用、分类构成和分组差异;不要输出用户级记录,不做运营决策。
tariff_describe will generate a set of files under outputDir, including data quality, numeric summaries, categorical distributions, binary feature summaries, grouped statistics, and a Markdown report.
Use Cases and Notes¶
- Use Cases: Suitable for scenarios requiring local tariff data processing, standardized field definitions, generating descriptive statistical reports, and ensuring data reproducibility.
- Notes:
- This plugin does not automatically make business decisions, and does not provide causal inference or recommendations for user operations actions.
- You must trust the source code to allow the pnpm prepare script to run.
- The data security boundary is to only read relative paths within the specified
dataDir. - Data identifiers are used only for deduplication and quality validation; no user-level records are returned.
After the steps above, the data preparation and descriptive statistics workflows are encapsulated as standardized tool calls. This plugin ensures the reproducibility of analysis results by locking environment dependencies and performing hash validation. For the related directory and source code, see: Plugin Directory | GitHub Repository.