Introduction

DeepSeek Harness (DSH) adopts a plugin-based architecture and supports connecting different LLM services through adapters. Kimi (Moonshot AI) provides a Code-specific API and a general open platform API, with differences between domestic and overseas regions. The dsh-llm-kimi plugin provides an LLM adaptation layer for Kimi models in DSH. It uses three routes to cover three Key types, addressing the integration challenges of multiple API endpoints and key management.

Core Features

This plugin implements the following capabilities:
* Multi-route adaptation: Supports both Kimi Code subscription and Moonshot open platform Keys, covering domestic and overseas regions.
* Thinking mode: Controls thinking depth via the reasoning_effort parameter.
* Tool calling and multimodal: Supports Function Calling and image input (via base64 data URI).
* Hot updates: Configuration and key changes take effect without restarting the process.

Installation and Enablement

Install the plugin from the command line. Build artifacts are committed to the repository, so no local compilation is required.

# 从 GitHub 安装
dsh plugin --profile web add github:haveanote06/dsh-llm-kimi

# 或从 npm 安装(发布后)
dsh plugin --profile web add dsh-llm-kimi

After installation, restart the dsh web service and configure the plugin on the “Kimi” settings page in the web UI.

Routes and Key Configuration

The plugin defines three routes, each corresponding to a different API endpoint and Key reference. When configuring, specify the route type and the corresponding API Key.

Route Details

Route Key Type Default Endpoint Default Key Environment Variable
kimi-code Kimi Code subscription (sk-kimi-*) https://api.kimi.com/coding/v1 KIMI_API_KEY
kimi-cn Moonshot open platform (domestic) https://api.moonshot.cn/v1 MOONSHOT_API_KEY
kimi-global Moonshot open platform (overseas) https://api.moonshot.ai/v1 MOONSHOT_API_KEY

Key Configuration Methods

Three configuration methods are supported; choose any one of them. By default, routes read the corresponding environment variable. They can also be overridden via apiKeyEnvCode/apiKeyEnvCn/apiKeyEnvGlobal or the shared apiKeyEnv.

  1. Environment Variable
    Set before starting dsh web:
   export KIMI_API_KEY=sk-kimi-...
   # 或
   export MOONSHOT_API_KEY=sk-...
  1. Web Credential Store (Recommended)
    Write directly to the credential store via the API, without restarting. After the write is complete, the corresponding row on the Models page displays a green “Configured” dot.
   curl -X POST http://127.0.0.1:3080/api/credentials.set \
     -H "Content-Type: application/json" \
     -d '{"type":"client-request","rpcId":"set1","method":"credentials.set","payload":{"ref":"KIMI_API_KEY","value":"<你的key>"}}'
  1. Credential File
    Append the value to ~/.dsh/.credentials.yaml (file permissions recommended: 0600):
   KIMI_API_KEY: <key>

Note: Due to a current DSH platform limitation, the built-in Models page renders namespaces other than llm-deepseek / llm-pi-ai (such as this plugin’s llm-kimi) as read-only. Therefore, Keys for third-party LLM plugins cannot be entered directly on the Models page; use one of the methods above or the built-in “Kimi” settings page of this plugin.

Models and Thinking Mode

Supported Models

The plugin provides a unified model catalog covering Kimi Code and open platform models. Context length and modalities are as follows:

Model Context Modality Notes
kimi-for-coding 256k Text+Image Kimi Code (K2.7 Coding)
kimi-for-coding-highspeed 256k Text+Image Kimi Code high-speed
k3 1M Text+Image Kimi Code K3
k3-256k 256k Text+Image Compact Kimi Code K3
kimi-k3 1M Text+Image Open platform K3
kimi-k2.6 / kimi-k2.7-code 256k Text+Image Open platform
moonshot-v1-8k/32k/128k 8k~128k Text Legacy open platform

Thinking Mode Parameter

Control thinking mode with the thinking parameter. The legacy enable_thinking boolean is silently ignored.

{
  "thinking": {
    "type": "enabled",
    "reasoning_effort": "low"
  }
}

reasoning_effort supports low, high, and max.

Protocol Obligations

The plugin strictly aligns with the official LLM adapter protocol:
* Streaming response order: The usage event precedes the finish event, and no content is emitted after finish.
* Tool parameters: Tool call parameters remain in RAW JSON format.
* Error handling: Follows dual-path error handling and respects options.signal for request cancellation.