AI Agent Hub
Back to skills
RAG Knowledge Curator icon

RAG Knowledge Curator

Knowledge Management Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_e9af5021/rag-knowledge-curator.

About this skill

Problem it addresses

Enterprise RAG systems often fail because retrieved chunks are noisy, fragmented, or hard to trace. The issue is frequently not the model, but the ingested corpus: navigation text, headers, repeated paragraphs, garbled characters, and ad-like noise can pollute vector indexes. Missing metadata such as topic, version, audience, confidentiality, and source makes retrieval results harder to explain. rag-knowledge-curator turns unstructured documents into a cleaner, auditable, RAG-ready dataset before they enter the vector store.

How it works

The skill runs a pre-ingestion curation workflow:
- Text cleaning: removes garbled text, headers, repeated paragraphs, ad-like noise, and invisible characters while preserving meaningful content.
- Chunking: splits text by chunk_strategy, keeping context boundaries intact and avoiding broken business logic, table explanations, or code blocks.
- Metadata extraction: tags chunks with topic, entities, version, audience, confidentiality, and source for filtering and traceability.
- Quality scoring: rates completeness, accuracy, timeliness, and readability, and notes why a chunk scores low.
- Versioned output: produces a summary, chunk previews, labels, scores, and curation recommendations that can feed vector database pipelines and human review.

Limits

It is best suited for curating documents, SOPs, technical manuals, and product guides. It does not replace vector database deployment, permission control, real-time data synchronization, or business review. For highly structured databases, log streams, or OCR output, it needs external tools for extraction and validation.

Use Cases

  • Clean and chunk technical manuals, then tag version, topic, and audience before RAG ingestion.
  • Review FAQ corpora before curation, flag low-quality chunks, and suggest completion or updates.
  • Convert product documentation into traceable chunks with quality scores and tags for vector pipelines.
  • Filter knowledge-base ingestion lists by topic, confidentiality, and source, then review low-scoring chunks.

Best For

  • Algorithm engineers building enterprise knowledge bases who need unstructured documents curated into retrievable chunks.
  • Operations or documentation teams maintaining SOPs and technical manuals who need cleaning, chunking, and confidentiality tagging before ingestion.
  • Data engineers building RAG pipelines who need ingestion lists with metadata and quality scores.
  • Knowledge review managers who need to inspect low-scoring chunks and decide on completion or updates.