Introduction

When developing agents with DeepSeek Harness (DSH), the web_search tool faces a classic tradeoff: fast but narrow queries versus slow but comprehensive queries. Hardcoding a single strategy means simple questions waste resources, while complex questions may receive insufficient information. The dsh-tool-bandit-search plugin uses a Thompson sampling contextual multi-armed bandit algorithm to let the tool automatically select search strategies based on actual usage.

Plugin Overview

This is a DeepSeek Harness plugin designed to replace the standard web_search tool. It introduces a new search tool that uses Thompson sampling to make adaptive decisions between two search strategies (quick and thorough). Maintainer: siruignaw-sys. License: MIT.

Core Features

  • Strategy Replacement: Replaces DSH’s built-in web_search with the search tool.
  • Dual-Strategy Selection:
  • quick: A single query, up to 5 results, fast, suitable for simple factual queries.
  • thorough: Runs three query variants (the original query + two rephrased perspectives) in parallel and merges them, up to 10 results, slower, suitable for open-ended questions.
  • Adaptive Learning: Uses the Thompson sampling algorithm, drawing from a Beta(α, β) distribution to select strategies.
  • Reward Calculation: After each invocation, computes a reward between 0 and 1 based on result count and response speed, then updates strategy parameters.

Installation and Activation

Use the following command to install the plugin:

dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-search

After installation, the Web UI must be restarted for the change to take effect. For local development, use dsh plugin --profile web add link:/absolute/path/to/dsh-tool-bandit-search.

# Restart Web UI
pnpm dsh web
# Or
dsh web

Runtime Example

During execution, the plugin outputs logs showing the selected strategy, the received reward, and the current Beta distribution parameters:

[bandit-search] arm=quick reward=1.000 durationMs=4393 resultCount=5 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1,"beta":1}}
[bandit-search] arm=thorough reward=0.854 durationMs=8481 resultCount=10 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1.85,"beta":1.15}}

Use Cases and Considerations

This plugin is suitable for scenarios where you want to dynamically optimize search strategies at runtime. Note the following:

  • State Storage: The bandit state is stored in memory and is reset after a process restart.
  • Reward Metric: The current reward is a heuristic metric (based on result count and latency) and does not represent answer relevance or correctness.
  • Optimization Scope: The plugin only optimizes strategy selection for a single invocation and does not control the model’s behavior across multiple search calls in one conversation.
  • API Stability: This plugin is built on the dsh developer preview, and the API may change.

Conclusion

dsh-tool-bandit-search provides a way for the DSH plugin ecosystem to automatically adjust tool behavior through online learning. It does not require an additional API key and directly reuses DSH’s ctx.web service. For more details and source code, refer to the catalog page or the GitHub repository.