RAG Knowledge Curator
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_e9af5021/rag-knowledge-curator.
About this skill
Problem it addresses
Enterprise RAG systems often fail because retrieved chunks are noisy, fragmented, or hard to trace. The issue is frequently not the model, but the ingested corpus: navigation text, headers, repeated paragraphs, garbled characters, and ad-like noise can pollute vector indexes. Missing metadata such as topic, version, audience, confidentiality, and source makes retrieval results harder to explain. rag-knowledge-curator turns unstructured documents into a cleaner, auditable, RAG-ready dataset before they enter the vector store.
How it works
The skill runs a pre-ingestion curation workflow:
- Text cleaning: removes garbled text, headers, repeated paragraphs, ad-like noise, and invisible characters while preserving meaningful content.
- Chunking: splits text by chunk_strategy, keeping context boundaries intact and avoiding broken business logic, table explanations, or code blocks.
- Metadata extraction: tags chunks with topic, entities, version, audience, confidentiality, and source for filtering and traceability.
- Quality scoring: rates completeness, accuracy, timeliness, and readability, and notes why a chunk scores low.
- Versioned output: produces a summary, chunk previews, labels, scores, and curation recommendations that can feed vector database pipelines and human review.
Limits
It is best suited for curating documents, SOPs, technical manuals, and product guides. It does not replace vector database deployment, permission control, real-time data synchronization, or business review. For highly structured databases, log streams, or OCR output, it needs external tools for extraction and validation.
Use Cases
- Clean and chunk technical manuals, then tag version, topic, and audience before RAG ingestion.
- Review FAQ corpora before curation, flag low-quality chunks, and suggest completion or updates.
- Convert product documentation into traceable chunks with quality scores and tags for vector pipelines.
- Filter knowledge-base ingestion lists by topic, confidentiality, and source, then review low-scoring chunks.
Best For
- Algorithm engineers building enterprise knowledge bases who need unstructured documents curated into retrievable chunks.
- Operations or documentation teams maintaining SOPs and technical manuals who need cleaning, chunking, and confidentiality tagging before ingestion.
- Data engineers building RAG pipelines who need ingestion lists with metadata and quality scores.
- Knowledge review managers who need to inspect low-scoring chunks and decide on completion or updates.
Related Skills
Search Huawei Cloud official docs and product pages to find ECS, OBS, RDS, CCE product specs, parameters, documentation, and API references without login.
OCR-based recognition for movie, train, flight, and event tickets in images or PDFs, extracting key fields into Markdown or JSON reports.
Turns notes, research, and meeting summaries into actionable next moves, plans, decisions, experiments, and decision-changing gaps.
A local wiki knowledge base manager that compiles raw documents into sourced, indexed Markdown pages with wikilinks, query support, and health checks.