Introduction

When performing model inference with DeepSeek Harness (DSH), if the underlying model service (for example, llama.cpp started with --parallel 1) cannot handle multiple requests simultaneously, excess requests are directly discarded by the server. The client cannot distinguish whether it is “waiting for a slot” or whether “the service is dead,” causing the Node HTTP layer to report terminated after 300 seconds. This issue frequently occurs when sub-agent and main-agent requests, or compaction and agent requests, overlap.

dsh-llm-gate is designed to solve the above problem. It maintains a FIFO queue inside the DSH process and temporarily stores excess requests until the model service frees up a slot before sending the HTTP request. This prevents requests from being discarded due to timeout while waiting and ensures that requests are processed in order.

Core Features

  • Per-provider concurrency control: Sets a concurrency limit for a specific model provider.
  • Request queue management: Temporarily stores requests exceeding the concurrency limit using a FIFO queue.
  • End-to-end coverage: Hooks the llm/stream waterfall and covers all model request scenarios, including agent, sub-agent, compaction, and title generation requests.
  • Error handling: Handles error states such as queue full (QUEUE_FULL) and timeout (QUEUE_TIMEOUT).

Installation and Activation

Use the official installation command to add the plugin to the specified configuration file (for example, web):

dsh plugin --profile web add dsh-llm-gate

After installation, you need to restart dsh web and open a new session to make the configuration effective. You can verify whether the configuration has been loaded correctly with the following command:

dsh --profile web --dump-config

Configuration Example

Add the configuration to the ~/.dsh/profiles/web/cordis.patch.yml file. The names corresponding to the providers key should match the route names in your llm-pi-ai.providers (or other adapter) settings. Providers not listed are not controlled by this plugin.

- id: llm-gate
  config:
    providers:
      llamacpp:
        maxConcurrent: 1
        maxQueued: 16
        queueTimeoutMs: 3600000

Configuration Options

Option Required Description
maxConcurrent Yes The number of requests allowed to be sent simultaneously to this provider. For llama.cpp, keep this value consistent with the --parallel parameter.
maxQueued No The number of requests allowed to wait in the queue. If this limit is exceeded, new requests fail immediately with error QUEUE_FULL. The default is unlimited.
queueTimeoutMs No The maximum duration a request can wait for a slot. If this time is exceeded, the request fails with error QUEUE_TIMEOUT. The default is to wait indefinitely.

Log Output

This plugin outputs the following logs to the DSH terminal to monitor request queueing and scheduling:

  • When queued: Shows that the current request has entered the queue and the queue depth.
    llm-gate: llamacpp session=a61e6e40 queued (depth 1)
  • When dispatched: Shows that the request was dispatched after waiting for a period of time.
    llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
  • Auxiliary requests: Requests with purpose=compaction or purpose=session-title are also logged.

Notes

  • Serialization limitation: This plugin serializes requests, so it does not make a single-slot server process faster. To improve performance, add slots to llama.cpp (for example, --parallel 2 --kv-unified) and increase maxConcurrent accordingly.
  • Timeout mechanism: Waiting time is not counted toward the adapter’s streamIdleTimeoutMs, because the adapter is not invoked before a slot is acquired. You still need to set a sufficient streamIdleTimeoutMs based on prompt processing time.
  • Request cancellation: Requests in the queue are canceled using an abort signal. If a stream is discarded without canceling it, the request remains in the queue until a slot is released, after which it is dispatched immediately and closed.
  • Dependency requirements: The plugin requires the llm service and hooks the llm/stream waterfall.

Ecosystem and License

The DSH plugin ecosystem emphasizes modularity and extensibility. As part of this ecosystem, dsh-llm-gate provides a controlled queue management mechanism for handling single-slot model services under high concurrency.