The DSH plugin yumu247/dsh-kb-rag provides DSH with a local-first RAG knowledge base toolkit. It builds on top of the kb-rag Python pipeline and supports semantic retrieval, incremental ingestion, web crawling, and related recommendations. This toolkit is designed to meet local deployment requirements, avoid cloud API calling costs, and ensure data remains on the local machine.

Core Features

The plugin wraps four core commands, each corresponding to a different knowledge base operation:

  • kb_query (semantic search): Performs semantic search in a local vector database and returns relevant text chunks along with their source and similarity scores.
  • kb_ingest (incremental ingestion): Imports a document directory into ChromaDB. Supports incremental updates based on manifest hashes and automatically handles SiYuan front matter metadata.
  • kb_crawl (batch web crawling): Uses Scrapling and markdownify to batch-crawl URLs and cleans the results into Markdown format.
  • kb_related (related recommendation): Recommends related documents for a known document based on retrieval or graph mode.

Installation and Prerequisites

Before installing the plugin, ensure that the host has a configured Python environment, Ollama models, and the kb-rag pipeline.

  1. Install the Python pipeline
    Clone the kb-rag repository and install the dependencies:
    git clone https://github.com/YuMu247/kb-rag
    cd kb-rag
    pip install -r requirements.txt
  1. Install the Ollama model
    Pull the bge-m3 embedding model:
    ollama pull bge-m3
  1. Configure the path
    Tell the plugin the installation path of kb-rag. This can be set through the plugin configuration or an environment variable:

    • Configuration method: Set kbRagDir in the DSH profile configuration.
    • Environment variable method: Set the KB_RAG_DIR environment variable to point to the kb-rag directory.
    • Default behavior: If not configured, the plugin looks for ./kb-rag under the session working directory.
  2. Install the DSH plugin
    Use the DSH plugin management command to install it:

    dsh plugin --profile <profile> add github:YuMu247/dsh-kb-rag
After installation, the four commands `kb_query`, `kb_ingest`, `kb_crawl`, and `kb_related` will be injected into the current DSH process. Before installing, confirm that the current DSH build supports the `dsh plugin` subcommand.

Typical Usage

kb_query

Used to query the local knowledge base.
* Parameters:
* query (required): Query text.
* k (optional): Number of results to return. Defaults to 5; maximum is 20.
* Return format: JSON format, containing matched text chunks and their sources.

kb_ingest

Used to import a document directory into the vector store.
* Parameters:
* docs: Defaults to <kb-rag>/docs.
* chunkSize: Chunk size. Defaults to 512.
* overlap: Chunk overlap size. Defaults to 64.
* force, reset, heading: Control overwriting and metadata handling.

kb_crawl

Used to batch-crawl web pages.
* Parameters:
* urls or urlsFile: Target URL list or file path.
* out: Output directory. Defaults to <kb-rag>/out.
* delay: Crawling interval in seconds. Defaults to 1.5.

Used for document related recommendation.
* Parameters:
* seed or fromDoc: Seed document ID (choose one).
* k: Number of items to return. Defaults to 5.
* graph: Graph mode switch. When greater than 0, enables graph-mode knowledge walking.

Notes

  • Local operation: The plugin runs entirely locally, requires no cloud API calls, and incurs zero cost.
  • Environment dependency: The plugin itself does not contain Python logic; it must be run with kbRagDir configured correctly.
  • Permissions and security: The plugin runs with the permissions of the current DSH process. Ensure that you have read and write permissions for the relevant files and directories.