Introduction

In the DeepSeek Harness (DSH) Web chat interface, when sending large requests, it is usually not immediately clear what the approximate cost of the call will be. The dsh-cost-estimate plugin inserts an inline notification in the chat stream before the response begins, estimating the input/output token range and cost range; after the response ends, the same line automatically updates to the actual usage, cost, and cache hit rate.

The difference from existing community plugins (such as dsh-session-cost) is that this plugin focuses on pre-estimation—because output length is inherently unpredictable, it provides ranges and continuously calibrates using actual historical values within the session.

Core Features

  1. Pre-estimation: Before the model starts generating, it first displays the estimated input tokens, output token range, and cost range.
  2. Post-response reconciliation: After the response ends (after receiving the provider’s actual usage), the same line switches to the actual tokens and cost, and shows the cache hit rate.
  3. Threshold triggering: By default, it displays only when estimated input >= 8000 tokens or estimated cost >= ¥0.01; small requests remain silent, while large requests are always flagged.
  4. Peak/off-peak pricing: Supports deepseek-v4-flash / v4-pro, automatically selecting the peak/off-peak pricing tier based on the current time.
  5. Input-side accuracy: Prioritizes anchoring to the actual usage returned by each request, using CJK-aware heuristic counting as an interim fallback.
  6. Output-side calibration: Uses a moving average to continuously refine subsequent estimates.
  7. Pure client-side: Calculation and rendering happen in the Web client, with zero intrusion into session logs.
  8. Multi-step support: Tool-call loops are annotated, and actual costs are accumulated.

Installation and Activation

dsh plugin --profile web add dsh-cost-estimate

After installation, restart dsh web and force-refresh the browser (Ctrl+Shift+R) to load the new plugin.

Notes:
* This command depends on pnpm (npm install -g pnpm).
* If Windows paths contain spaces, use a junction or a path without spaces.
* Estimates are for reference only; actual DeepSeek billing takes precedence.

Configuration

The following keys are supported under config in cordis.patch.yml:

  • minInputTokens: default 8000. Display only when the estimated input exceeds this token count.
  • minCostCny: default 0.01. Display only when the estimated cost exceeds this CNY amount.
  • defaultModel: default deepseek-v4-flash. The fallback model before the first usage anchor.
  • headerTokensEstimate: default 6000. The heuristic token count for the system prompt + tool schema before the first usage anchor.
  • defaultCacheHitRatio: default 0.5. The assumed cache hit rate before the first usage anchor.

Example configuration:

- id: cost-estimate
  name: dsh-cost-estimate
  config:
    minInputTokens: 20000
    minCostCny: 0.05

Effect Example

After sending a large question, before the response:

预估:输入约 9.9K tok · 输出 2.2K–7.3K tok · 费用约 ¥0.0096–¥0.0198(v4-flash)

After the response ends, the same line updates to:

实际:输入 9.1K tok · 输出 1.1K tok · 费用 ¥0.02(缓存命中 87% · v4-flash)

Notes

  • The input side is accurate after anchoring; the output side is a range estimate.
  • The plugin runs with the permissions of the current dsh process; check the source code and license before installing.

Summary

dsh-cost-estimate is suitable for developers who need fine-grained cost control and evaluation of large model call overhead during development and debugging. It implements real-time monitoring and estimation of DeepSeek API costs on the client side.

  • Plugin directory: https://www.skillhub.cn/plugins/Yvesgao/dsh-cost-estimate
  • Source repository: https://github.com/Yvesgao/dsh-cost-estimate