AI Agent Hub
Back to skills
Token Optimizer Plus Proxy icon

Token Optimizer Plus Proxy

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_884ff669/token-optimizer-plus.

About this skill

Problem: LLM calls waste tokens on repetition and noise

When AI agents call OpenAI, Anthropic, or Ollama frequently, prompts often contain repeated context, verbose explanations, tables, and long histories. Sending this directly can increase cost and push requests close to the context window, forcing manual truncation or loss of relevant details.

How it works: a local transparent proxy plus a four-step pipeline

token-optimizer-plus runs as a local proxy and, by default, listens on localhost:8899. After pointing the application API base URL to it, requests pass through an optimization pipeline before being forwarded upstream. The pipeline includes:
- Semantic deduplication: reduce near-duplicate context
- Exact deduplication: remove identical fragments
- Noise filtering: drop low-value text
- Semantic compression: optional further compression using local Ollama
It also maintains data/to_stats.db and shows cumulative saved tokens, cost, strategy contribution, session history, and request details at http://localhost:8899/dashboard, making it easier to identify which content consumes the most tokens.

Boundaries: best for text APIs, not every integration

This tool mainly optimizes text requests and does not process images or audio. By default it binds to localhost, forwards over HTTP, and has limited optimization for SSE / streaming responses. Before production use, run the dashboard self-check and back up the data/ directory regularly.

Use Cases

  • Debug high token usage in an agent app by routing its API through a local proxy and inspecting deduplication, filtering, and compression.
  • Control OpenAI or Ollama API spend by reviewing cumulative saved tokens, cost, and per-strategy contribution on the dashboard.
  • Keep long conversations within the context window by filtering noisy text and compressing repeated context before forwarding requests.
  • Trace wasted tokens in repeated code, tables, or verbose text by expanding request details and viewing the strategy breakdown.

Best For

  • Engineers maintaining AI agent applications who want to reduce per-call LLM token cost without changing business code.
  • Developers using OpenAI or Ollama who want to quantify how much request text is wasted by repetition, noise, or tables.
  • Technical leads owning model budgets who need cumulative savings, session history, and strategy contribution to assess cost.
  • Application developers working on context engineering who need automatic compression of long chats or documents to reduce overflow risk.