AI Agent Hub
Back to skills
Local RAG Builder icon

Local RAG Builder

AI Agent Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_800d68d6/local-rag-builder according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

Local RAG setups often stall on environment setup, model downloads, chunking quality, and retrieval noise: missing Python packages, failed embedding model fetches, split code blocks and tables, and multiple knowledge bases bleeding into each other. local-rag-builder turns these steps into a debuggable pipeline so an agent gets usable context before answering.

How It Works

  • Setup: rag_env_setup.py checks Python 3.11+, installs chromadb, sentence-transformers, and langchain dependencies.
  • Model and ingestion: downloads embedding models from multiple sources with retries and path correction; documents are chunked by text_splitter.py, stored in Chroma, and deduplicated by SM3 hash.
  • Chunking and retrieval: fixed window, recursive, header hierarchy, sentence, and semantic chunking; GuardStack protects code, formulas, tables, and HTML; optional reranking and knowledge-base routing are available.
  • Two modes: rag_skill.py retrieves only and lets the agent answer; rag_standalone.py retrieves, then calls LM Studio, Ollama, or vLLM to generate the answer.

Boundaries

Best for local private docs, Markdown/PDF material, and single-process knowledge bases. Native text formats work best; PDF/image OCR must be enabled; it does not directly support OpenAI/Cohere API embeddings; keep each base under about 50k entries and avoid multi-user concurrent writes.

Use Cases

  • Ingest internal runbooks from Markdown and PDF, apply header-based chunking, and answer troubleshooting queries.
  • When embedding model downloads fail, retry via ModelScope or HuggingFace mirrors and validate the local cache path.
  • Keep project documents in separate knowledge bases, auto-classify imports, and retrieve similar documents for queries.
  • Preserve code blocks and tables during context building by configuring GuardStack and post-processing sub-chunking.

Best For

  • Engineers maintaining private documentation who want local retrieval without uploading source files.
  • AI app developers needing retrievable context for agents without deploying another LLM.
  • Technical writers handling Markdown, PDF, and structured docs who need controllable chunking strategies.
  • Local inference users running LM Studio, Ollama, or vLLM who want retrieval plus generation in one flow.