AI Agent Hub
Back to plugins
🤖

dsh-ctx-probe

Model Inference Updated 2026.09.08

Run the following command in DeepSeek Harness:

dsh plugin install IYIcode/dsh-ctx-probe

Paste the following prompt into your AI chat to install this plugin:

Run dsh plugin install IYIcode/dsh-ctx-probe in DeepSeek Harness to install this plugin; the source is available at https://github.com/IYIcode/dsh-ctx-probe .

About this plugin

When you run a local llama.cpp or Ollama model through DSH, llm-pi-ai falls back to a 262144-token contextWindow for any model that does not declare one. If your server was actually started with -c 65536, the built-in 80 percent pressure compaction can never reach its trigger, and the first oversized request hard-fails with a 400 overflow. Manually editing the setting is easy to forget, and a server restart with a different window size forces another round of edits. dsh-ctx-probe was built for exactly this: before every model call it queries the server for the real runtime context window (meta.n_ctx, not the training-time n_ctx_train) and aligns the DSH contextWindow to that value. Shrink the server and the plugin tightens before overflow; grow the server and it widens automatically. Both directions, zero numbers to type.

The probe strategy covers four endpoints, namely llama.cpp /v1/models, /slots, /props, and Ollama /api/show, adopting the first usable result. Loopback addresses are re-probed on every request, so a server restart is picked up on the very next message; remote addresses use a roughly ten-minute TTL cache with single-flight and exponential backoff to avoid probe storms. All errors are swallowed on the plugin side, so a failed probe can never break the agent loop.

If you are wiring local inference servers into DSH and would rather not keep tuning the context window by hand, this plugin keeps working after a one-time Web Host restart. It is fully self-contained, does not touch node_modules, depends on no @deepseek-ai package, and differs from manual-catalog or scheduled-sync alternatives by probing per request and syncing both ways.

Use Cases

  • Local llama.cpp started with a smaller -c flag than the DSH contextWindow assumes, causing 400 overflow
  • Ollama models without a declared contextWindow where the built-in fallback is too large for compaction to trigger
  • Server restarts with a changed window size and the next model call must reflect the new value immediately

Best For

  • Developers wiring local inference servers into DSH who want to stop hand-tuning the context window
  • Users deploying local LLMs via llama.cpp or Ollama and frequently changing the -c flag
  • DSH users who need probe failures to never break the agent loop