DeepSeek Harness (DSH) uses a plugin-based architecture, and its extensibility is implemented through plugins. In research and analysis scenarios, manually collecting information across platforms, cross-validating it, and generating reports is inefficient. dsh-harvest is a native multi-platform research pipeline plugin that provides end-to-end automation from discovery, scraping, verification, and audit. It runs on Windows, macOS, and Linux.

The plugin is maintained by user toustifer under the MIT open source license. It provides five native tools that form a scout → extract → verify → audit data flow, outputting structured JSON and text for agents to read.

Core Capabilities

The plugin provides the following five native tools. Each tool has a clear failure strategy to ensure that a single point of failure does not affect the overall pipeline:

  1. harvest_scout: Discovers candidate sources in parallel across multiple platforms. It supports 11 data channels: GitHub, Web, Twitter, Reddit, Xiaohongshu, YouTube, Bilibili, V2EX, LinkedIn, Telegram, and LINUX DO. If a single channel fails (for example, the corresponding CLI is not installed or the user is not logged in), it logs [SKIP] and continues running without blocking subsequent steps.
  2. harvest_deep_research: Asynchronous end-to-end deep research. It is used to generate long-form comprehensive research reports and sources, with a timeout/exception-safe fallback mechanism.
  3. harvest_extract: Extracts body content one by one. It supports RSS parsing and Xiaoyuzhou podcast transcription (requires a groq key), with automatic escalation of fetching strategies (such as Jina Reader) and paywall/unreachable tagging.
  4. harvest_verify: Cross-source verification of claims. It compares consistency for key claims and outputs one of three states: consistent, weak, or unverified.
  5. harvest_audit: Five-dimensional source credibility audit. It scores sources and outputs trusted, cautious, or discard states.

Data Channels and Dependencies

The plugin connects to platforms through underlying CLI tools. If the corresponding CLI is not installed, the channel is skipped automatically.

Required CLI Tools

  • gh: GitHub search.
  • mcporter: Web/Exa search and LinkedIn job posting scraping.
  • opencli: Scraping social platforms such as Twitter, Reddit, and Xiaohongshu.
  • yt-dlp: YouTube video scraping.
  • bili-cli: Bilibili video scraping (note: this is a third-party unofficial client; obtain it from trusted sources).
  • mcp-telegram: Telegram search (requires a configured Python environment or specified path).
  • linuxdo-mcp: LINUX DO search (requires a configured cookie).
  • feedparser: RSS parsing (a Python package).

Platform Differences and Notes

  • Windows: The Xiaoyuzhou channel is unavailable; mcporter / opencli must be invoked via PowerShell shims.
  • macOS / Linux: If DSH is launched from Finder or a LaunchAgent, PATH issues may occur (for example, tools under /opt/homebrew/bin cannot be found). Launch it from a terminal or explicitly configure PATH.
  • Bilibili: It depends on the unofficial bili-cli client, which may involve compliance and security boundary risks; verify it yourself.

Installation and Enablement

It is recommended to install via the DSH native marketplace or CLI command.

  1. Install the plugin:
    dsh plugin --profile web add @stifer/dsh-harvest
  1. Configure and enable:
    Append the following to <dshHome>/profiles/web/cordis.patch.yml:
    - insert:
        - id: harvest
          name: '@stifer/dsh-harvest'

Typical Usage

The following are examples of the main interface calls provided by the plugin:

  • Broadcast discovery across all platforms:
    harvest_scout(query="最新 AI 项目", limit=8)
  • Targeted channel discovery:
    harvest_scout(query="DeepSeek 优化实践", platforms=["github", "linuxdo", "telegram", "web"], limit=5)
  • End-to-end deep research:
    harvest_deep_research(input="深入剖析大模型推理与架构演进")
  • Deep content extraction:
    harvest_extract(urls=["https://…", "https://…"])
  • Cross-source verification:
    harvest_verify(claims=["DeepSeek 出了视觉模型"], sources=[…])
  • Source credibility audit:
    harvest_audit(sources=[{title, url, type}])

Configuration

The plugin supports configuring services such as Tavily via environment variables or DSH’s native credential management.

1. Deep Intelligence Channel Configuration

Telegram and LINUX DO require configuring the Python interpreter path and session information:

  • TELEGRAM_PYTHON: Python interpreter path for Telegram.
  • TELEGRAM_MCP_PATH: Project directory for mcp-telegram.
  • LINUXDO_PYTHON: Python interpreter path for LINUX DO.
  • LINUXDO_MCP_PATH: Project directory for linuxdo-mcp.
  • LINUXDO_COOKIE: Session cookie for LINUX DO. Skipped automatically if not provided.

2. Tavily Deep Research Configuration

Supports DSH credential management or system environment variables. Configure in ~/.dsh/settings.yaml:

harvest:
  tavilyApiKeyEnv: TAVILY_API_KEY
  tavilyEndpoint: https://search.cliproxyapi.xyz/search
  tavilyResearchEndpoint: https://search.cliproxyapi.xyz/research

Then save the key in ~/.dsh/.credentials.yaml.

Summary

By standardizing multi-platform research workflows, dsh-harvest reduces the complexity of agent-driven research. It supports stable operation in mixed environments through a modular toolset and graceful degradation strategies. Plugin directory page: https://dshfind.com/zh/plugins/toustifer/dsh-harvest
Code repository: https://github.com/toustifer/dsh-harvest