Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Coordinate hierarchical coding-agent fleets for repository-wide audits, implementation sprints, migrations, and complex work that benefits from parallel specialists. Use when a user asks for subagents, a fleet, parallel delegation, GPT-5.6 Sol, Terra, or Luna routing, broad codebase completion, or independent implementation and verification passes. Enforce bounded ownership, concurrency-aware waves, dirty-worktree safety, runtime-honest model handling, and evidence-based integration.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 119% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 89% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 485% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -33% | 0% |
Coordinate specialists while retaining responsibility for integration and the final result. Use native collaboration tools; do not build an orchestration framework inside the target repository unless the user explicitly requests one.
Treat findings as inputs to action, not the end product. When the user authorizes implementation, carry every confirmed in-scope finding through disposition, build, integration, and verification. Do not stop after producing an audit report.
gpt-5.6-sol for ambiguous integration and hard engineering, gpt-5.6-terra for exploration and bounded implementation, and gpt-5.6-luna for high-volume mechanical verification.gpt-engineer skill is installed, use its guarded scripts/run_codex_agent.py fallback for explicitly model-pinned delegates. Otherwise do not spawn a generic, inherited, model-less, GPT-5, GPT-5.4, Spark, or Claude fallback; keep work with the exact-model parent or report the blocker.Before delegation, read the applicable AGENTS.md files and capture a baseline with git status --short, the current branch, the complete dirty-path ledger, and relevant diffs. Start metadata-first: do not dump secret-bearing files, huge binaries, generated artifacts, untracked directories, or submodule contents into context. Treat every pre-existing change as user-owned.
Keep one orchestrator responsible for decomposition, status, integration, and final verification. Use specialists directly for a small fleet. With four or fewer total slots, keep spawning centralized under the orchestrator; allow a lead to spawn children only when the task is larger, ownership remains explicit, and capacity has been reserved.
For a repository-wide completion pass, prefer this first wave:
| Profile | Mode | Responsibility | | --- | --- | --- | | Sol / sol_engineer | Bounded writes when authorized | Resolve ambiguous architecture, hard implementation, integration, and root-cause debugging | | Terra / terra_explorer | Read-only | Trace architecture, SDK usage, dependencies, documentation, and incomplete product paths | | Terra / terra_worker | Bounded writes | Implement isolated, well-specified findings with focused tests | | Luna / luna_verifier | Verification | Run test matrices, diff hygiene, residual scans, and acceptance evidence | | Orchestrator | Integrator | Resolve overlaps, own cross-cutting judgment, and decide final acceptance |
When Codex custom profiles are present and selectable, route by their exact agent types. A profile file alone is not proof that the child used its model. Keep final semantic acceptance and authority decisions with the orchestrator even when Luna performs mechanical verification.
Give every spawned agent a bounded contract containing:
Use prompts that are independently actionable. Do not ask multiple writers to fix anything they find across the same repository.
Choose fast for one known path, standard for two or more independent shards, and broad only for repository-scale uncertainty. A fast task may stay with the orchestrator and focused checks; it does not require an inventory fleet.
TODO string with a defect.After the first cycle, work delta-only: reuse accepted evidence, inspect confirmed residuals and changed paths, and rerun only invalidated checks. Do not restart broad discovery or full verification without new cross-cutting evidence.
Maintain a finding ledger across the waves. Give each finding an identifier, evidence, severity, affected paths, dependencies, owner, planned verification, and one final disposition: implemented, already satisfied, invalid, duplicate, blocked, or explicitly deferred. Never silently drop a finding between research and build.
Do not spawn more workers than the runtime supports. Use later waves or reuse an idle agent when specialist continuity is useful.
Inspect agent state after spawning and at wave boundaries. Integrate communications deliberately:
Never treat an agent's completion message as proof by itself. Inspect its artifacts and rerun proportionate checks from the orchestrator context.
Fleet completion includes teardown. Before returning control:
A finished subagent does not imply that its MCP servers or runtime helpers were reclaimed. Native runtime helpers may be shared or retained by the host, so never terminate the Codex app, shared MCP services, another task's cohort, or an unclassified process. If the runtime owns a residual helper and exposes no safe task-scoped teardown, report it precisely instead of using a broad process kill.
Define completion from the user's acceptance criteria when available. Otherwise label each candidate as a confirmed defect, probable gap, informational cleanup, or unverified suspicion before assigning work.
Require the narrowest relevant tests after each write scope, then broader integration checks. Inspect every command first for network access, code generation, database connections, filesystem mutations, and production defaults. Typical gates include:
git status and diff inspection against the captured baseline.Classify results as passed, failed, or not run with a reason. Do not call a feature complete because its stub was removed; prove its user-facing path, error behavior, persistence or integration boundary, and regression coverage where applicable.
Inspect verification scripts before running them. If a smoke, integration, or release command can default to production, require an explicit non-production target or skip it and report why.
Keep audits and local implementation inside the authorized repository. Do not deploy, mutate production data, send messages, push, merge, or install global tools unless the user separately authorized that action. Prefer deterministic validation over a forward test that could touch production.
When current best practices matter, verify unstable claims with official primary sources. Adapt guidance to the repository's actual stack instead of forcing fashionable migrations.
Lead with the outcome. Include:
Claim complete only when the requested scope and verification gates are satisfied. Otherwise state precisely what remains partial.
Other measured skills in the registry, with their headline benchmark lift.