GPT Image 2 Multi-Model Image Generation
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md and install @user_fdb96f95/gpt-image-2-1.
About this skill
Problem
AI agents often lack a stable image-generation entry point. API keys may come from environment variables or host configuration, while endpoints, models, sizes, and response formats vary. Hard-coding one provider makes the workflow brittle across gateways, compatible APIs, or local models. This skill turns image generation into a reusable agent capability with a consistent discovery and execution path.
How It Works
- API discovery: reads environment variables first, then host files such as
config.yamlorsettings.jsonto find an image endpoint. - Model selection: prefers
gpt-image-2, thendall-e-3,dalle-3,stable-diffusion-xl, or any available model, with endpoint validation. - Structured prompts: assembles
Subject / Action / Context / Composition / Lighting / Styleto reduce unstable outputs from vague prompts. - Generation and saving: calls
/v1/images/generations, handlesb64_jsonorurlresponses, and writes images to a local directory. - Recovery and memory: maps errors such as
401,403,429, andtimeout, retries with exponential backoff, and updates session notes.
Boundaries
It works best when an OpenAI-compatible image API is already configured and the agent needs to pick models automatically. It will still fail if no usable API exists or the endpoint does not support image generation. Keep model size limits in mind: gpt-image-2 supports 1024x1024 and 2048x1024, while dall-e-3 commonly supports 1024x1024 and 1792x1024.
Use Cases
- Generate images in an agent workflow by calling configured GPT Image 2, DALL·E 3, or open-source models in priority order.
- When a host exposes only an OpenAI-compatible endpoint, discover the API, validate models, and save b64_json or URL images.
- Compose prompts for detailed scenes by organizing subject, action, context, composition, lighting, and style.
- Handle 429 or timeout errors with exponential backoff, retry, model switching, and operational notes.
Best For
- Engineering teams that need agents to select usable image models and save results locally
- Agent developers maintaining OpenAI-compatible gateways who need one image-generation entry point
- Design-collaboration engineers creating concept art or visual assets with structured prompt templates
- Backend engineers handling API rate limits, timeouts, or endpoint differences who need stable retries
Related Skills
Uses step-by-step choices to confirm business, palette, and layout, then exports an editable .drawio architecture diagram.
An agent skill for the Miaoyin AI music REST API that supports SUNO/Mureka-based song generation, continuation, cover, video, WAV, and stem splitting.
Distills text, files, or images into structured knowledge and generates knowledge-card prompts for drawing tools, supporting 3:4 portrait or 16:9 landscape outputs.
Capture full-page, viewport, or element screenshots with Playwright while handling lazy loading, internal scroll, mobile layouts, and Linux CJK fonts.