Introduction¶
DeepSeek Harness (DSH) adopts a plugin-based architecture, allowing developers to extend Web GUI capabilities through configuration files. In use, developers need to view different models’ performance on benchmarks, trend changes, and costs within the same interface. The dsh-models-radar plugin connects directly to public benchmark data sources and integrates model capability radar charts into DSH’s settings and conversation flow, solving the problem of fragmented multi-model comparison.
What Is This¶
This is a plugin that provides model capability radar display for the DeepSeek Harness Web GUI, maintained by hi-fangj. It reads public benchmark data from deng.codexradar.com, adds a Model Radar page to Settings through DSH configuration slots, and displays real-time benchmark scores for the current session model next to the model selector in the conversation interface.
Core Features¶
-
Settings Integration
- Adds a Model Radar page to Settings through thesettings.sectionslot.
- Adds a Model Radar card under Plugins -> Configurable plugins through thesettings.plugin.itemslot. -
Capability Overview
- Groups and displays by base model, supporting expansion to view different reasoning-effort levels.
- Uses a fixed 0–110 IQ scale and displays Harness attribution badges on each row (such as Codex, Claude Code, DSH, etc.). -
Trends and Ratings
- Provides trend tabs for the last 24 hours and the last 7 days, each with independent scaling.
- Includes community rating cards, supporting 0–10 experience scores over 7-day/24-hour windows. -
Cost and Metrics
- Includes a Cost × IQ card with tabs for overall cost, time cost, and price cost.
- Displays efficiency metric badges: IQ, average cost, average duration, cache hit rate, and 24-hour run count. -
Benchmark Channels
- Supports two channels:deep-swe(code repair tasks, binary-choice voting) andpompeii-adjacency(visual reconstruction tasks, continuous F1 scores). -
Conversation UI Interaction
- Displays the current session model’s real-time DeepSWE score on the left side of the tool bar in the conversation editor.
- Clicking the score opens a capability overlay to view detailed cross-model comparison data. -
Lightweight and Privacy
- The browser does not request upstream directly; requests are forwarded via a same-origin host proxy. Within the data freshness window, only expired data is requested, while the rest uses local cache.
- No credentials are required; when offline, it falls back to the latest persisted snapshot.
Installation and Enablement¶
Run the following command in the terminal to install the plugin into the web configuration profile:
dsh plugin --profile web add github:hi-fangj/dsh-models-radar
The plugin repository already includes built Host and browser packages, so installation can be done directly without running dependency scripts. After installation, refresh http://127.0.0.1:3080. If the DSH process has not hot-loaded, it must be restarted once.
Typical Usage¶
- Open the DeepSeek Harness Web GUI, go to Settings -> Model Radar, and view the capability overview grouped by base model.
- Switch models in the conversation editor; the current score for that model appears on the left side of the tool bar (formatted as
model@reasoningEffort). - Click the score to open the capability overlay and view comparison details for different base models as well as 24h/7d trends.
- Switch tabs in the Cost × IQ card to evaluate model cost-effectiveness by time or price dimensions.
Applicable Scenarios and Notes¶
Suitable for developers who need to directly compare model capabilities side by side, account for costs, and monitor trends within the DSH interface. The plugin depends on DeepSeek Harness Web GUI, dsh-super-injector (runtime or persistent local installation), and Node.js 22 or later.
Note that the plugin runs with the permissions of the current DSH process. It is recommended to review the source code and license before installing. The plugin does not request or submit any credentials.
Summary¶
This plugin integrates fragmented benchmark data, trend charts, and cost analysis into DSH’s settings and conversation flow. Its lightweight design provides offline fault tolerance without credentials, making it easy to quickly evaluate model performance in a local environment.