Introduction¶
When performing model inference with DeepSeek Harness (DSH), if the underlying model service (for example, llama.cpp started with --parallel 1) cannot handle multiple requests simultaneously, excess requests are directly discarded by the server. The client cannot distinguish whether it is “waiting for a slot” or whether “the service is dead,” causing the Node HTTP layer to report terminated after 300 seconds. This issue frequently occurs when sub-agent and main-agent requests, or compaction and agent requests, overlap.
dsh-llm-gate is designed to solve the above problem. It maintains a FIFO queue inside the DSH process and temporarily stores excess requests until the model service frees up a slot before sending the HTTP request. This prevents requests from being discarded due to timeout while waiting and ensures that requests are processed in order.
Core Features¶
- Per-provider concurrency control: Sets a concurrency limit for a specific model provider.
- Request queue management: Temporarily stores requests exceeding the concurrency limit using a FIFO queue.
- End-to-end coverage: Hooks the
llm/streamwaterfall and covers all model request scenarios, including agent, sub-agent, compaction, and title generation requests. - Error handling: Handles error states such as queue full (
QUEUE_FULL) and timeout (QUEUE_TIMEOUT).
Installation and Activation¶
Use the official installation command to add the plugin to the specified configuration file (for example, web):
dsh plugin --profile web add dsh-llm-gate
After installation, you need to restart dsh web and open a new session to make the configuration effective. You can verify whether the configuration has been loaded correctly with the following command:
dsh --profile web --dump-config
Configuration Example¶
Add the configuration to the ~/.dsh/profiles/web/cordis.patch.yml file. The names corresponding to the providers key should match the route names in your llm-pi-ai.providers (or other adapter) settings. Providers not listed are not controlled by this plugin.
- id: llm-gate
config:
providers:
llamacpp:
maxConcurrent: 1
maxQueued: 16
queueTimeoutMs: 3600000
Configuration Options¶
| Option | Required | Description |
|---|---|---|
maxConcurrent |
Yes | The number of requests allowed to be sent simultaneously to this provider. For llama.cpp, keep this value consistent with the --parallel parameter. |
maxQueued |
No | The number of requests allowed to wait in the queue. If this limit is exceeded, new requests fail immediately with error QUEUE_FULL. The default is unlimited. |
queueTimeoutMs |
No | The maximum duration a request can wait for a slot. If this time is exceeded, the request fails with error QUEUE_TIMEOUT. The default is to wait indefinitely. |
Log Output¶
This plugin outputs the following logs to the DSH terminal to monitor request queueing and scheduling:
- When queued: Shows that the current request has entered the queue and the queue depth.
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
- When dispatched: Shows that the request was dispatched after waiting for a period of time.
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
- Auxiliary requests: Requests with
purpose=compactionorpurpose=session-titleare also logged.
Notes¶
- Serialization limitation: This plugin serializes requests, so it does not make a single-slot server process faster. To improve performance, add slots to
llama.cpp(for example,--parallel 2 --kv-unified) and increasemaxConcurrentaccordingly. - Timeout mechanism: Waiting time is not counted toward the adapter’s
streamIdleTimeoutMs, because the adapter is not invoked before a slot is acquired. You still need to set a sufficientstreamIdleTimeoutMsbased on prompt processing time. - Request cancellation: Requests in the queue are canceled using an abort signal. If a stream is discarded without canceling it, the request remains in the queue until a slot is released, after which it is dispatched immediately and closed.
- Dependency requirements: The plugin requires the
llmservice and hooks thellm/streamwaterfall.
Ecosystem and License¶
- License: MIT
- Repository: https://github.com/d3vmeh/dsh-llm-gate
- Directory page: https://www.skillhub.cn/plugins/d3vmeh/dsh-llm-gate
The DSH plugin ecosystem emphasizes modularity and extensibility. As part of this ecosystem, dsh-llm-gate provides a controlled queue management mechanism for handling single-slot model services under high concurrency.