DeepSeek Harness (DSH) adopts a plugin-based architecture. When building multi-agent systems, the main agent typically calls the expensive official API, while subagents, if not configured, continue to use the official interface, increasing costs and creating privacy leakage risks. dsh-llama-responses solves this problem as an adapter plugin; it routes subagent requests to a locally running llama.cpp service through the OpenAI Responses protocol.

Plugin Overview

  • Name: dsh-llama-responses
  • Owner: SnowRikka
  • Category: Model Inference
  • License: MIT
  • Core Features:
    • Multi-turn memory
    • function_call streaming
    • function_call_output return
    • Subagent delegation (spawn/continuable background, fork/one-shot foreground)
    • Follow-up queuing (send_message)
    • Status query (list_agents)
    • Interrupt control (interrupt_agent)
    • Settlement notification (report)
    • Skill discovery and loading

Installation

Install using the official plugin command:

dsh plugin --profile web add git+https://github.com/SnowRikka/dsh-llama-responses.git

Before npm release, it can also be installed via symbolic link (for source modification):

git clone https://github.com/SnowRikka/dsh-llama-responses.git ~/dsh-llama-responses
ln -sfn ~/dsh-llama-responses ~/.dsh/profiles/web/node_modules/dsh-llama-responses

Prerequisites

  1. Local model service is running: Ensure llama.cpp’s llama-server is running and listening on a local port (such as 8080).
  2. Critical configuration step: After each DeepSeek Harness startup, you must first manually select the local qwen3.8-27b once and send one prompt, then switch the main agent back to the DeepSeek V4 model. Without this step, subagents will not be routed to the local model correctly.

Enable Subagent Features

This plugin is mounted on DSH’s subagent system. By default, the subagent tools in the web profile are disabled and need to be enabled manually.

Add the following configuration to ~/.dsh/profiles/web/cordis.patch.yml:

- id: tool-subagent-control        # send_message / interrupt_agent
  disabled: false

- id: tool-subagent-list-agents    # list_agents
  disabled: false

- id: tool-subagent                # 委派工具(可继续后台子代理)
  disabled: false

- id: tool-subagent-fork           # 委派工具(one-shot,已重定向 spawn=无父对话)
  disabled: false

After enabling, the model will have access to five tools: subagent, subagent_fork, send_message, interrupt_agent, and list_agents.

Enable Local Model Routing

1. Configure Adapter

Register the adapter in ~/.dsh/profiles/<profile>/cordis.patch.yml:

- insert:
    - id: llm-responses-local
      name: dsh-llama-responses
      config:
        baseURL: http://127.0.0.1:8080/v1
        providerName: llama-responses
        apiKeyEnv: DSH_LOCAL_LLM_KEY
        models:
          - id: qwen3.8-27b
            name: Qwen3.8-27B IQ4_XS (local)
            contextWindow: 163840
        defaultContextWindow: 163840
        defaultMaxTokens: 8192

2. Configure Subagents to Use the Local Model

Specify the subagent’s provider as llama-responses in the patch:

- id: tool-subagent
  config:
    provider: spawn
    toolName: subagent
    backgroundMode: continuable
    agentOptions:
      provider: llama-responses
      model: qwen3.8-27b
      maxTokens: 8192

3. Environment Variable

Set the local model authentication placeholder:

export DSH_LOCAL_LLM_KEY=local-no-auth

Subagent Isolation Design

This is an important design feature of this plugin. All subagents do not inherit the parent conversation, including completed history content from the parent session. The subagent_fork tool has been redirected via a patch to ensure complete isolation. If the parent agent needs to provide background information, it must explicitly include it in the delegation instructions.

Protocol Details

The plugin implements adaptation for the OpenAI Responses protocol. When integrating with llama.cpp, pay attention to the following details:

  • History message format: History message entries must include the top-level type: 'message'; assistant content uses output_text.
  • Tool calls: Tool definitions use the flat format {type: 'function', name, description, parameters}.
  • Result return: Tool result return uses {type: 'function_call_output', call_id, output}.
  • SSE events: Listen for response.output_text.delta, response.function_call_arguments.delta, response.output_item.done, and response.completed.