AI Agent Hub
Back to skills
LLM Benchmarking and Self-Optimization icon

LLM Benchmarking and Self-Optimization

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_fcba917f/llm-benchmark.

About this skill

Problem

Teams comparing LLM API performance often face unstable metrics, missing first-token latency, token throughput, percentile stats, and no durable stress-testing workflow. Manual scripts also tend to miss dimensions such as temperature, max_tokens, concurrency, and multilingual behavior, making results hard to reproduce and archive.

How It Works

This skill turns LLM API testing into a repeatable workflow:
- Benchmarking: run quick or full modes to collect TTFT, TPS, ITL, P50/P90/P99, and success rate.
- Stress testing: specify duration and concurrency, with checkpointing, exponential backoff, and crash recovery.
- Reporting and comparison: generate HTML/PDF reports, compare models side by side, and infer trends from historical runs.
- Self-optimization: use optimize suggest and auto-tune to recommend or grid-search parameters such as temperature and max_tokens, with scoring targets like tps, ttft, and balance.

Scope and Caveats

The API endpoint should be OpenAI-compatible, such as /v1/chat/completions. Auto-tuning sends real requests, so it can affect quotas and runtime; stress tests should be run in a controlled environment. Local Ollama can be tested through the compatible endpoint, usually with the key set to ollama.

Use Cases

  • When onboarding an OpenAI-compatible API, verify first-token latency, throughput, and success rate against release requirements.
  • Before launch, run continuous load tests for local Ollama and cloud models to confirm stability and checkpoint recovery.
  • Compare candidate models by TPS, TTFT, and composite score, then export an HTML report for model-selection review.
  • Recommend max_tokens and temperature for production or development scenarios, and grid-search the optimal combination.

Best For

  • Engineers handling model integration: they need a fast check of API connectivity, latency, and throughput before release.
  • Product or algorithm leads selecting models: they need side-by-side comparisons and exportable review reports.
  • SRE or platform engineers running load tests: they need long-duration stability, checkpoint recovery, and trend analysis.
  • Prompt or service-parameter engineers: they need recommendations for temperature and max_tokens based on historical results.