LLM Benchmarking and Self-Optimization
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @user_fcba917f/llm-benchmark.
About this skill
Problem
Teams comparing LLM API performance often face unstable metrics, missing first-token latency, token throughput, percentile stats, and no durable stress-testing workflow. Manual scripts also tend to miss dimensions such as temperature, max_tokens, concurrency, and multilingual behavior, making results hard to reproduce and archive.
How It Works
This skill turns LLM API testing into a repeatable workflow:
- Benchmarking: run quick or full modes to collect TTFT, TPS, ITL, P50/P90/P99, and success rate.
- Stress testing: specify duration and concurrency, with checkpointing, exponential backoff, and crash recovery.
- Reporting and comparison: generate HTML/PDF reports, compare models side by side, and infer trends from historical runs.
- Self-optimization: use optimize suggest and auto-tune to recommend or grid-search parameters such as temperature and max_tokens, with scoring targets like tps, ttft, and balance.
Scope and Caveats
The API endpoint should be OpenAI-compatible, such as /v1/chat/completions. Auto-tuning sends real requests, so it can affect quotas and runtime; stress tests should be run in a controlled environment. Local Ollama can be tested through the compatible endpoint, usually with the key set to ollama.
Use Cases
- When onboarding an OpenAI-compatible API, verify first-token latency, throughput, and success rate against release requirements.
- Before launch, run continuous load tests for local Ollama and cloud models to confirm stability and checkpoint recovery.
- Compare candidate models by TPS, TTFT, and composite score, then export an HTML report for model-selection review.
- Recommend max_tokens and temperature for production or development scenarios, and grid-search the optimal combination.
Best For
- Engineers handling model integration: they need a fast check of API connectivity, latency, and throughput before release.
- Product or algorithm leads selecting models: they need side-by-side comparisons and exportable review reports.
- SRE or platform engineers running load tests: they need long-duration stability, checkpoint recovery, and trend analysis.
- Prompt or service-parameter engineers: they need recommendations for temperature and max_tokens based on historical results.
Related Skills
Provides Claw with character-library selection, switching, saving, and global SOUL.md style sync for role-based conversation.
An AIONE Agentic AI Infrastructure SDK wrapper for building production AI agents with memory, skills, workflows, and hooks.
A Python/TypeScript SDK wrapper for the DeepSeek-Reasonix native AI coding agent, with prefix-cache support.
A browser automation tool for analysts, operators, and developers that locates elements, fills forms, extracts structured content, and supports no-code scheduling and export.