Introduction

When processing local model inference, DeepSeek Harness (DSH) may encounter issues where the model thinks for too long or generates large files (such as write tool calls), causing connection interruptions. This happens because the Node.js built-in fetch (used by DSH model adapters) times out if the server has not sent response headers or transferred any data bytes within 300 seconds. DSH itself has no configuration options for these two timings, so requests are terminated after the default 5 minutes. The dsh-fetch-timeouts plugin solves this by raising the HTTP timeout thresholds for the Node process.

What Is This

  • Plugin name: dsh-fetch-timeouts
  • Maintainer: d3vmeh
  • Core value: Raises Node.js HTTP timeout settings to prevent slow local models from being cut off after 5 minutes of silence.
  • Problem solved: When a server (such as Ollama or LM Studio) does not send response headers or has no data transfer for a long time, Node’s default timeout mechanism terminates the request.
  • License: MIT

Core Features

  • Raises Node.js HTTP timeout settings (global effect).
  • Replaces Node’s global fetch dispatcher with an undici Agent.
  • Supports HTTP_PROXY and HTTPS_PROXY environment variables.
  • Can be tuned through the cordis.patch.yml configuration file.

Installation and Activation

To install the plugin, use the official command:

dsh plugin --profile web add dsh-fetch-timeouts

After installation, the timeout is raised to 30 minutes by default. To customize it, edit the configuration file ~/.dsh/profiles/web/cordis.patch.yml:

- id: fetch-timeouts
  config:
    headersTimeoutMs: 3600000   # 允許在收到響應頭前的等待時間;0 表示禁用
    bodyTimeoutMs: 3600000      # 允許在接收數據塊之間的等待時間;0 表示禁用

After configuring, restart dsh web. A confirmation message will appear in the startup log:

fetch-timeouts: headers 1800000 ms, body 1800000 ms (process-wide)

Typical Usage

In addition to configuring the plugin timeouts, you also need to configure watchdogs in the provider route, otherwise the watchdogs will trigger before the fetch timeout. Add the following to the corresponding provider configuration:

llm-pi-ai:
  providers:
    ollama:
      streamIdleTimeoutMs: 1800000
      timeoutMs: 1800000

Applicable Scenarios and Considerations

  • Applicable scenarios: Used with local model servers such as Ollama or LM Studio, where the server does not send keepalive pings (heartbeat packets) during long inference or large file generation. The llama-server from llama.cpp sends a ping every 30 seconds by default, so llama.cpp users usually do not need this plugin.
  • Global impact: The plugin effect is process-level. Any fetch call using the Node global dispatcher (including model calls, web searches, HTTP MCP servers, and cloud providers) is affected. web_fetch is not affected because it uses its own per-request agent.
  • Dead connection detection: Because timeouts are extended, truly dead connections take longer to detect.
  • Version requirements: Requires Node.js >= 22.19.0 and undici >= 8.
  • Uninstall behavior: After uninstalling the plugin, it reverts to the default undici dispatcher, not the original Node.js fetch implementation.
  • Long-term plan: This is a stopgap solution until DSH natively exposes these timeout settings (the pi-ai dependency already accepts a custom fetch).

Brief Conclusion

The dsh-fetch-timeouts plugin provides a necessary timeout buffer for long-running local model inference by modifying the global fetch dispatcher. In shared-host environments or scenarios requiring strict timeout control, evaluate its global effect carefully. Plugin source code and detailed documentation can be found in its GitHub repository.