Preface¶
When developing with DeepSeek Harness (DSH), developers often need to monitor the model’s inference process and performance metrics. Although the native DSH interface provides basic information, during streaming output, real-time data about reasoning progress, first token latency, and token generation rate per second are often not intuitive enough.
The dsh-response-meta plugin is designed to address this pain point by embedding a lightweight runtime summary directly in the response stream of the Web interface, enabling developers to monitor the model’s status without leaving the interface.
Plugin Positioning¶
dsh-response-meta is a DSH Web plugin. It adds a real-time runtime summary to each model response, including the model name, reasoning level, tokens per second, timestamp, total runtime, and first token latency. This summary is displayed immediately when the model responds and updates as the data stream progresses; it does not wait for the response to complete.
Core Features¶
- Real-time summary updates: During streaming output, it displays the model, reasoning level, and estimated throughput (tokens/s) in real time.
- Complete metrics display: After the response is complete, it shows the model name, tokens/s, timestamp, total runtime, and TTFT (time to first token).
- UI optimization: It automatically hides duplicate native assistant timestamps while preserving user message timestamps and the native usage details button.
- Interruption support: For manually interrupted responses, including interruptions during the reasoning phase, the summary is also displayed in the message stream.
Installation and Configuration¶
The plugin requires DSH version 0.1.5-rc.1 and the Web profile. Version v0.4.0 optimizes handling of transient events and packaged streams.
- Run the installation command:
dsh plugin --profile web add github:Unintendedz/dsh-response-meta#v0.4.0
- After installation is complete, the running DSH Web service must be restarted to load the new plugin.
Usage Examples¶
The plugin displays different summary content in the response stream depending on the status:
- During streaming output:
deepseek-v4-pro · thinking 1.5k chars · ~42 tok/s - Response completed:
deepseek-v4-pro · thinking 15 tok · 46 tok/s · 14:16 · runtime 2m24s · first token 2.1s - Interrupted response:
deepseek-v4-pro · thinking 1.5k chars
Notes¶
- Version dependency: DSH version 0.1.5-rc.1 or higher is required.
- Rendering logic: Missing fields are omitted. If all fields are unavailable, nothing is rendered.
- Model inheritance: In subsequent rounds, if no new request header is present, the most recently known model name is inherited.