Oracle AI/ML Design & Evaluation Expert
Paste the following prompt into your AI chat to install this skill:
Follow https://skillhub.cn/install/skillhub.md to install @org-02qudk26/oracle-zh.
About this skill
Problem
LLM features often get stuck between prompts, RAG retrieval, agent tool use, safety guardrails, and cost/latency budgets. Oracle is an AI/ML design and evaluation skill that turns AI requirements into testable, implementable, and monitorable specs. It does not write production code directly; it focuses on system prompts, RAG, LLM-as-judge, OWASP LLM Top 10, and token budgets.
How It Works
Oracle follows ASSESS → DESIGN → EVALUATE → SPECIFY:
- Assess the current prompt versioning, retrieval quality, guardrails, evaluation sets, and cost/latency baselines, then flag anti-patterns and blocking gates.
- Design prompt patterns, RAG chunking/hybrid search, agent/tool contracts, model routing, and graceful degradation paths.
- Evaluate with golden test sets, Recall@5, Faithfulness, regression gates, and LLM-as-judge calibration to avoid single-judge bias.
- Hand off implementation contracts to Builder, test strategy to Radar, security requirements to Sentinel, RAG ingestion to Stream, and monitoring SLOs to Beacon.
It treats prompts as versioned code, prioritizes retrieval quality over model size, and designs safety as layered guardrails. Every proposal should include token budgets, p95 latency, and release gates.
Boundaries
Use it for AI feature design, prompt/RAG/agent architecture, LLM safety, evaluation frameworks, and cost/latency optimization. Route to Builder, Gateway, Sentinel/Probe, Radar, or Nexus when the main task is implementation, general API contracts, penetration testing, test automation, or multi-agent orchestration. Avoid shipping unevaluated prompts, using unvalidated LLM output in critical decisions, and hard-coding model names, and consider RAG poisoning, position bias, length bias, and cache cost tradeoffs.
Use Cases
- Design a RAG support pipeline in a SaaS product, defining chunking, hybrid search, and Recall@5 gates.
- Turn prompt drafts into versioned specs with golden test sets, regression gates, and LLM-as-judge calibration.
- Define agent tool contracts and MCP descriptions with fallback paths, token budgets, and safety guardrails.
- Assess cost and latency before AI release, setting p95, cache hit, and safety trigger thresholds.
Best For
- Backend engineers owning AI feature architecture need to turn retrieval, prompts, and evaluation into implementation contracts.
- Product engineers building RAG support or knowledge assistants need chunking, retrieval SLOs, and release gates.
- Platform engineers focused on LLM safety and compliance need to encode OWASP risks and PII policies into specs.
- SREs managing AI cost and latency need model routing, caching, and p95 alert thresholds.
Related Skills
Model routing, persistent parameter management, self-check repair, and global default model control for XiaoYi Claw.
Install, update, and manage OpenClaw Skills through SkillHub, with automatic detection after installation.
Provides API endpoints for AI agents to post bottles and graffiti, browse the feed, and interact with likes and comments.
ReqPlan constrains agent development, debugging, and analysis workflows with a seven-stage state machine, checkpoints, quality audits, and local Harness artifacts.