Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn an existing audit, finding list, issue set, review, failing-test report, or implementation plan into completed, verified code through a coordinated agent fleet. Use when the user asks to build from findings, implement every audit item, finish a known backlog, remediate review results, or continue from research without repeating the whole investigation. Validate each finding, preserve repository state, assign non-overlapping writers, integrate in dependency order, and track every item to an
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 243% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-12 | ✓→✗ | ▼ Worse | 422% | 0% |
Convert findings into working software. Own the implementation result; do not merely redistribute the list or return another plan.
Trace each finding to current repository truth. Mark it confirmed, already satisfied, duplicate, invalid, or blocked. Do not implement stale advice blindly. Research only the gaps needed to make a safe build decision; avoid restarting a broad audit unless the findings are unusable.
If current APIs, dependencies, standards, or security guidance affect the implementation, verify them with primary sources. Adapt the result to the repository's actual stack.
Order confirmed work by dependency and blast radius:
Use the runtime's available concurrency, counting the orchestrator as a slot. Give one writer ownership of each file or tightly coupled subsystem. Latest-only is the default: use confirmed sol_engineer (gpt-5.6-sol) for hard integration, terra_worker (gpt-5.6-terra) for bounded implementation, terra_explorer for read-heavy gaps, and luna_verifier (gpt-5.6-luna) for mechanical checks. If the native spawn schema cannot prove an exact route and the sibling gpt-engineer skill is installed, use its guarded CLI runner. Otherwise never use a generic, inherited, model-less, GPT-5, GPT-5.4, Spark, or Claude substitute; keep the work with the exact-model parent or report the limitation.
Every writer contract must include exact paths, success criteria, prohibited side effects, required tests, baseline constraints, and expected handoff evidence. Keep overlapping work read-only.
For each wave:
Do not drop difficult items. A confirmed in-scope finding must end as implemented or blocked by a concrete missing authority, credential, external dependency, or mutually exclusive user decision. Do not use deferred unless the user explicitly accepts deferral.
After all waves, run checks selected by the changed behavior and acceptance criteria:
Keep credentialed, destructive, deployment, messaging, merge, and push actions outside scope unless the user separately authorized them.
After accepting the last implementation result, wait for every required handoff and inspect the live agent tree. Interrupt superseded or stale agents, then verify that no bounded worker remains active.
Track every background process or temporary resource launched by the fleet. Preserve evidence, then stop and reap task-owned subprocess groups, watchers, servers, and listeners and remove task-owned temporary worktrees unless the user explicitly requested continued runtime. Classify ownership using agent state, parent process, working directory, launch time, and recorded PID or resource identifier. Never kill by process name alone, and never terminate shared MCP services, the host application, another task's cohort, or an unclassified process. Report host-retained helpers when the runtime provides no safe task-scoped teardown.
Return the outcome first, followed by the finding ledger disposition summary, changed subsystems, verification matrix, fleet teardown result, preserved user work, and exact blockers. Claim completion only when every confirmed finding has an acceptance result, no required build work remains, and task-owned resources have been reclaimed or explicitly retained.
Other measured skills in the registry, with their headline benchmark lift.