Preface¶
DeepSeek Harness (DSH) supports plugin extensibility. This plugin targets the notice and announcement scenarios at the China Agricultural University (CAU) and provides an open-source crawler and AI processing pipeline. It integrates public web scraping, AI summary generation, classification, and deadline extraction, displays results in the DSH sidebar panel, and supports direct querying in conversations.
Unlike hosted notification aggregation services, this project open-sources only the tool itself. Users collect and store the data themselves in their own private repositories; the data is not part of the tool service. Maintainer: @ZBber-lab.
Core Capabilities¶
The plugin mainly addresses the issues of scattered information and time-consuming reading, providing the following features:
- Sidebar Panel: Requires DSH version ≥ 0.1.5-rc.2. The panel docks to the official DSH right sidebar and includes Today’s Briefing, Key News, My Items (deadline reminders), category channels, follow and archive, data management, and article reading (supports original text + AI summary + deadline highlighting + one-click citation into the conversation).
- Category Channel Management: In settings, you can toggle whether to display homepage sources (by default, new sources are displayed; turning it off only hides them, while data continues to be crawled). Supports “Add New Source”, using
tools/scraper/probe.mjsto probe site structures and write the confirmed configuration. - Multi-System Parsing: Built-in parsers for Boda (boda), Sudy (sudy), and CAU News Network custom (news-custom), driven by
sites.json.probe.mjscan probe target sites and generate configuration snippets, making it convenient to integrate other universities or department/faculty channels. - AI Processing: Uses the DeepSeek API to perform article summarization, classification (notices/news/lectures, etc.), importance assessment, and deadline extraction.
- Conversation Querying (MCP): Provides 6 MCP tools (latest notices, keyword search, deadline items, site directory, usage statistics, article details), allowing the AI to answer queries directly in conversations.
- Data Autonomy: The data repository is fully configured by the user (owner/repo); tokens are stored locally only. The plugin and MCP share the same credential configuration.
Installation and Enabling¶
Installing the plugin requires a DSH environment.
Desktop Version (Recommended):
Click “Plugins” → “Add Plugin” in the DSH sidebar, enter the following address, and install:
github:ZBber-lab/cau-portal-open
After installation, be sure to click “Enable Now” in the plugin list. A “CAU Portal” entry will appear at the bottom of the sidebar; if not enabled, the entry will be grayed out.
CLI / Web Environment:
If your environment has the dsh command, use the following commands (note the profile):
dsh plugin --profile desktop add "github:ZBber-lab/cau-portal-open"
# 或
dsh plugin --profile web add "github:ZBber-lab/cau-portal-open"
After successful installation, you need to reload DSH for changes to take effect.
Data Pipeline and Usage¶
The plugin data flow is: crawler fetches public web pages → AI processing → artifacts are committed to the user’s private data repository → the plugin and MCP read and display them.
Local Crawling¶
The tool code is located in the tools/scraper/ directory. Before running, prepare two repositories: one is the plugin repository (to obtain tool code), and the other is your data repository (to store crawl results).
- Crawling:
node tools/scraper/crawl.mjs --data-dir ../<你的数据仓>/data --pages 2 --articles 8
- AI Processing:
DEEPSEEK_API_KEY=sk-xxx node tools/scraper/enrich.mjs --data-dir ../<你的数据仓>/data --limit 8
Artifacts will be written to index.json, feed/, articles/, and usage.jsonl under the <data repository>/data/ directory. After completion, you need to commit and push in the data repository; otherwise, the plugin panel will not refresh.
Scheduled Automatic Crawl (Optional)¶
To avoid manual runs, you can use GitHub Actions. Copy .github/workflows/crawl.yml from the tool repository to your data repository.
Since GitHub free private repositories do not support native schedule triggers, you typically need an external scheduling service (such as cron-job.org) to call the Actions workflow_dispatch endpoint to perform automatic crawling every 2 hours.
MCP Conversation Integration¶
To enable the AI to call tools directly in conversations, an MCP client must be registered.
- Install dependencies:
cd tools/mcp && pnpm install
- Configure
cordis.patch.yml:
Add the following to the corresponding DSH profile configuration file (path depends on environment; desktop version usually~/.dsh/profiles/desktop/cordis.patch.yml):
- id: mcp-cau
name: '@deepseek-ai/dsh-mcp-client'
serverName: cau
transport: stdio
command: node
args: [<本仓库路径>\tools\mcp\index.mjs]
cwd: <本仓库路径>\tools\mcp
*Note: You do not need to duplicate the data repository configuration in this file (since v0.5.4); the plugin automatically reads the configuration from the panel settings page.*
- Configure Token:
In the DSH sidebar “CAU Portal” → Settings → Token Management, paste your GitHub data repository fine-grained token.
Prerequisites and Notes¶
- Environment Dependencies:
- DSH official desktop version (recommended) or Web/CLI environment, version ≥ 0.1.5-rc.2 (this version is required for docked right sidebar panel).
- Node.js 18+.
- Data Source:
- A GitHub private repository is required as the data repository.
- A DeepSeek API Key is required (for pipeline AI processing).
- Security:
- GitHub tokens are stored only on local machines; do not leak them.
- The cron-job.org token for scheduling should only be granted
Actions: Read & writepermissions. - Campus data copyright belongs to the university; publicly posted page information is for personal learning use only.
Conclusion¶
The CAU Portal plugin encapsulates crawler technology and AI capabilities as part of the DSH ecosystem. Through the panel and MCP tools, it makes searching and reading notices more efficient and controllable. All data is self-hosted by users, and the tool code is open source and transparent. Project directory and GitHub repository: zbber-lab/cau-portal-open.