Preface¶
DeepSeek Harness (DSH) is used to debug and evaluate Agent behavior, but in long execution chains, developers find it hard to intuitively see the Agent’s actual flow over time, failure nodes, and Token consumption. dsh-maze draws the Agent’s execution process as a “maze map,” placing the main path, failure branches, and backtracking points on the same timeline, and provides data tracks and an analysis panel.
What is this¶
dsh-maze is an execution maze plugin for DeepSeek Harness, maintained by user lamost423.
The plugin was originally named dsh-trace-compare and was renamed to dsh-maze starting with v1.0.0. The old package remains installable but is no longer updated. Migration only requires two commands (see the migration notes for the specific commands).
It mainly addresses visualization problems in Agent debugging: through a timeline, data tracks, and deterministic analysis, it lets developers clearly see how the Agent reaches its result step by step, rather than relying on the LLM’s self-report.
Core features¶
Maze timeline
Plots the main path, failure branches, and backtracking points of the execution process on the same timeline. The system automatically collapses idle intervals longer than 60 seconds (shown as a ⏸ annotation) and aggregates dense operation segments into “×N” badges (click to zoom in and view label details). The progress bar uses failure heat colors, so even an 8-hour long session remains fully legible.
Data tracks
Below the timeline, it provides data tracks synchronized with the maze:
* Tool call density: Colored by tool type, showing a tick line for each call.
* Token pulses: A stacked bar chart displaying Token counts for cached input, uncached input, reasoning, and visible output.
* Context pressure: A line + area chart showing model window usage percentage and 70%/90% threshold lines. Context compression events are annotated as “⌄−N%”; hover to view the ground-truth values before and after compression.
Deterministic execution analysis
All statistics are deterministic aggregations over decision data and do not invoke the LLM.
* Failure recovery chain: Records how each failed call is handled afterward (retry unchanged / change parameters / switch tool / not recovered) and the elapsed time.
* Execution analysis area: Provides a duration distribution scatter plot, a tool outcome matrix (including success rate, P50/P95, and maximum duration), and task completion assessment (task completed / tests / build / Lint / artifacts / manual confirmation).
Multi-session comparison
Supports same-axis comparison of 2–5 runs of the same task on different models. Features include round alignment lines (automatically connecting the answer node of each round), manual anchors (pinning semantically equivalent moments), and branch inventory (listing the branch step count and duration delta for each lane by round).
Replay and export
Supports replaying the entire execution process at up to 300× speed. Supports one-click export of the current view (including zoom window and filter state) as SVG or 2x PNG. Exports use a fixed light background and are suitable for sharing.
Other features
* Bilingual UI: Supports switching between Chinese and English.
* Theme following: Automatically follows the host DSH light/dark theme.
* Agent relationship graph: Visualizes a star-shaped overview of the main Agent and subagents.
Typical usage¶
Live maze
In the DSH conversation tab, the maze grows in real time with the current session. The “Execution Maze” entry in the sidebar can also open the live maze tab directly.
View or compare
Click the “Execution Maze” item in the sidebar, then upload a session log file to view historical sessions or compare them.
Locate and trace
Click any node or arc on the maze map, and a detail panel opens on the right, showing the full command, returned content (first 5,000 characters), duration, and judgment basis. The detail panel provides “Locate this step in conversation,” which jumps back to the conversation page and highlights the corresponding line.
Export view
The maze map can be exported as an SVG or 2x PNG image.
Notes¶
- Data integrity: The plugin’s aggregation over decision data is deterministic and does not invoke the LLM. Missing data is honestly labeled (tracks without data are not drawn, and unknown model windows are not guessed); nothing is ever fabricated.
- Legacy version notice: The old
dsh-trace-comparepackage remains installable but is no longer updated. - Dependency on host capabilities: The subagent branch feature in the live tab depends on the host’s ability to load child session history in the background.
Conclusion¶
dsh-maze uses visualization to integrate the Agent execution maze, data tracks, and deterministic analysis into a single interface. For developers who need to deeply understand Agent logic, troubleshoot long execution chains, or compare the performance of different models, this is a practical debugging tool.