Introduction

As session data accumulates, determining which tasks a model handles stably and which areas are prone to failure often requires manually reviewing large volumes of logs. dsh-cap-profile converts local session history into visualized capability profiles and error-rate dashboards, directly answering the question: “Which model is stable on which tasks, and which tasks does it often fail on?”

What Is It

This is a plugin for DeepSeek Harness (DSH), maintained by Ansonfishing. It converts session history from ~/.dsh/sessions into tool-usage and error-rate dashboards, supports filtering by model and time range, and provides a multi-model comparison view.

Installation and Enablement

Prerequisites: DSH (web required) + pnpm; Node ≥ 24 (for zstd decoding).

1、Switch to the DSH web profile directory:

cd ~/.dsh/profiles/web

2、Install the plugin:

pnpm add github:Ansonfishing/dsh-cap-profile

3、Configure and restart:
Add "dsh-cap-profile" to the dsh.profile.bundles array in package.json, then restart dsh.

After enabling, a “Capability Profile” tab appears in the session view. The first open triggers a background initial scan (approximately 20 seconds for large histories); mock data is displayed in the meantime.

Core Features

Model Comparison Table
Displays each model’s (provider / model) session count, tool call count, error count, and error rate.

Tool Top 5 / Error Signature Top 5
Lists each model’s most commonly used tools and highest-frequency error signatures. Tool error rates are marked by color:
* 0: gray
* <5%: amber
* ≥5%: red

Error signatures are usually normalized high-frequency patterns, such as bash: [exit code: 137] OOM, used to quickly identify where a model is likely to fail.

Multi-Model Comparison (≤4)
Select multiple models to enter the comparison view. The view displays each model’s data as cards:
* Convergence metrics comparison (best/worst/top highlighted)
* Top-10 tool matrix
* Top 5 frequent errors
* Daily calls/errors trend mini bar chart (current day’s errors marked in red)

If fewer than 2 valid models are selected, it automatically falls back to the single-model detail view.

Time Range Filtering
Supports all / last 7 / last 30 / last 90 days / today / yesterday. The anchor is the maximum date in the data, making it robust against system clock drift.

Security and Performance
* Read-only: analyzes and displays only, and never modifies session data. Routing validates client headers + Origin/Referer; cross-origin requests return 403.
* Incremental cache: based on per-file mtime baselines, it performs background incremental updates every 60s and full refreshes every 24h. While the initial scan is incomplete, the route does not block and automatically falls back to the previous cache or mock data.
* Zero runtime dependencies: pure Node.js implementation.

Typical Usage

1、Filter the time range: Select “last 30 days” in the time filter to observe a model’s performance fluctuations over the past month.
2、Compare multiple models: Select 2–4 models (for example, GPT-4o and DeepSeek-V3) to enter the comparison view and compare convergence metric differences in tool calls and error rates.
3、Locate failure points: View “Error Signature Top 5” on the model detail page. If a model frequently shows bash: [exit code: 137] OOM, you can adjust that model’s resource limits or prompts accordingly.

Notes

The plugin runs with the permissions of the current dsh process, and the analysis process only reads ~/.dsh/sessions. To debug the panel in an environment without DSH, you can clone the repository and open test/harness/index.html directly in a browser, switching data scenarios via ?scenario=mock|live|empty|error.

References