Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Autonomously pursue a sustained engineering outcome through repeated research, planning, multi-agent implementation, integration, verification, and gap-closing cycles. Use when the user asks for a /goal-style run, says do not stop, finish the whole codebase, research and build autonomously, babysit an outcome, or wants the orchestrator to keep working across continuations until genuinely complete. Use native goal tracking when explicitly requested and available; otherwise maintain an equivalent
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 753% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 41% | 0% |
Own a durable engineering outcome from uncertainty to verified completion. Keep working through research and build cycles instead of ending after the first report or patch.
do not stop expands persistence, not authority.If the runtime exposes native goal tools and the user explicitly requested a goal or autonomous sustained outcome, create or adopt one concrete objective. Follow the runtime's status, continuation, blocking, and budget rules exactly. Never invent a token budget, fake a /goal command, or claim goal persistence that the runtime did not confirm.
Otherwise maintain a goal ledger in the working plan containing:
Keep the ledger stable across continuations and context compaction. Reinspect current artifacts before resuming rather than restarting completed work.
Run one complete cycle, then repeat only while a confirmed gap remains:
Do not stop after research when the objective includes building. Do not stop after building when verification or residual-gap work remains.
The first cycle may be broad. Every later cycle is delta-only: retain the goal and finding ledgers, inspect only confirmed residuals and changed paths, and rerun only verification invalidated by new integration. Never restart repository-wide research just because automatic continuation is active.
Inspect the live agent tree and tool schema before choosing a topology. Count the orchestrator as a concurrency slot. Use later waves rather than oversubscribing the runtime.
sol_engineer / gpt-5.6-sol routing for ambiguous architecture, hard implementation, and integration.terra_explorer and terra_worker / gpt-5.6-terra routing for read-heavy research and bounded routine implementation.luna_verifier / gpt-5.6-luna routing for high-volume mechanical verification and residual scans.gpt-engineer skill is installed, use its guarded Codex CLI fallback. Otherwise never use generic, inherited, model-less, GPT-5, GPT-5.4, Spark, or Claude substitutes; keep work with the exact-model parent or record the blocker.At every cycle boundary, reconcile the finding ledger with both the live agent tree and a task-owned resource ledger. Do not start a new wave while superseded workers or their owned subprocesses are still consuming capacity.
When a worker finishes, collect its handoff, verify its artifacts, wait for its terminal state, and reclaim its task-owned subprocess groups, watchers, listeners, temporary worktrees, and other temporary resources. Preserve evidence first. Use recorded ownership plus parentage, working directory, launch time, and agent state; never kill by executable name alone. Shared MCP services, the Codex host, another task's cohort, and unclassified processes are outside the cleanup envelope.
Before marking the goal complete, interrupt stale or invalidated agents, inspect the live tree again, and compare the final process and listener inventory with the baseline. If the host retains a runtime helper and offers no safe task-scoped teardown, record that residual explicitly rather than broad-killing it.
Choose the next action by leverage: unblock critical dependencies, close user-facing paths, remove false completion signals, and strengthen verification before cosmetic cleanup. When an approach fails, diagnose it, try safe alternatives, and record the evidence.
Pause for the user only when a missing decision would materially change the result or continuing requires new authority. Treat missing credentials or external state as a reported verification boundary, not permission to fabricate success.
Use the runtime's native blocked status only under its stated threshold and semantics. Difficulty, uncertainty, slow progress, or a nearly exhausted budget are not blockers by themselves.
Do not install a generic Stop hook to force persistence. Such hooks can create unbounded continuation loops and cannot determine whether new user authority is required. Prefer native goal state or the explicit goal ledger above.
Finish only when:
Mark a native goal complete only after those conditions hold. Return the outcome, cycles completed, finding dispositions, verification matrix, fleet teardown result, external-only checks, residual risks, and exact next action if anything remains.
Other measured skills in the registry, with their headline benchmark lift.