AI Agent Hub
Back to skills
Auto Agent Long-Task Engineering Framework icon

Auto Agent Long-Task Engineering Framework

AI Agent Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please follow https://skillhub.cn/install/skillhub.md to install @user_21e7e3b2/auto-agent-skill.

About this skill

Problem Being Addressed

Long coding tasks often span multiple files and are prone to context pollution: a model that implements and reviews its own work can mistake assumptions for facts and miss its own errors. Auto Agent reduces this by decomposing work into small, verifiable tasks and separating roles.

How It Works

  • Role separation: Planner breaks down tasks and analyzes architecture; Generator implements and self-tests; Evaluator independently reviews code against the contract.
  • Sprint Contract: each task writes sprint-contract.md with goals, acceptance criteria, target files, definition of done, and test method.
  • Task decomposition: each atomic task touches 1-3 files, and tasks.json tracks dependencies, acceptance_criteria, steps, and complexity.
  • Evaluation matrix: scores functional correctness, architecture compliance, code quality, and reusability; any hard threshold failure sends the work back for fixes.
  • Session recovery: progress.md, handoff.md, and code files restore context; three failed iterations mark the task blocked.

Boundaries and Notes

Use it for large refactors, new modules, new systems, or explicit long-running work; avoid single-file fixes, queries, and simple bugs. It expects separate git commit for each task and human acceptance by default at Level 1. Failures must not be reported as passing.

Use Cases

  • When adding a new module across multiple files, use it to draft sprint contracts, atomic tasks, and acceptance criteria before implementation.
  • Before a large refactor, decompose requirements into atomic tasks limited to 1-3 files and track dependencies, acceptance criteria, and target files.
  • After implementation, have an independent Evaluator score the code against the sprint contract and reject it below hard thresholds for fixes.
  • After a session interruption, restore task status from tasks.json, progress.md, and handoff.md instead of guessing the prior context repeatedly.

Best For

  • Backend engineers landing large refactors or new systems who need complex requirements split into verifiable, contract-based atomic tasks
  • Architects maintaining shared codebases who want an independent Evaluator to review Generator output against architecture rules
  • Developers using Claude Code for long tasks who want Sprint Contracts to constrain implementation and acceptance
  • Engineering leads enforcing quality gates who require below-threshold work to be rejected and fixed before merge