Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Autonomous iteration loop: modify, verify, keep/discard against any metric
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 125% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 577% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 1% | 0% |
Iterations: unlimited.autoresearch/{subcommand}-{YYMMDD}-{HHMM}/ directory.handoff.json. Evals reads *-results.tsv./autoresearch)Parse the invocation in this order:
| Condition | Mode | |---|---| | Metric: or Verify: present | Classic — existing metric loop, unchanged | | Free-form natural-language goal, no metric/verify | Orchestrator — see Orchestrator section | | Nothing | Setup wizard — interactive config builder | | --classic flag | Force Classic regardless of goal text | | --auto flag | Force Orchestrator regardless of goal text |
Print a banner on every invocation: [autoresearch] mode: classic | orchestrator | wizard.
| Command | Does | Default Iterations | |---|---|---| | /autoresearch | Iterate against a metric: modify → verify → keep/discard | 25 | | /autoresearch_plan | Convert a goal into validated Scope, Metric, Verify config | N/A | | /autoresearch_debug | Hunt bugs: hypothesize → test → falsify → repeat | 15 | | /autoresearch_fix | Crush errors one-by-one until zero remain | 20 | | /autoresearch_security | STRIDE + OWASP audit with red-team personas | 15 | | /autoresearch_ship | Ship through 8 phases: checklist → dry-run → deploy → verify | N/A | | /autoresearch_scenario | Generate edge cases across 12 dimensions | 20 | | /autoresearch_predict | 5 expert personas debate before implementation | N/A | | /autoresearch_learn | Scout codebase → generate docs or wiki → validate → fix loop | 10 | | /autoresearch_reason | Adversarial debate with blind judges until convergence | 8 | | /autoresearch_probe | 8 personas interrogate requirements until saturation | 15 | | /autoresearch_improve | Research ICP challenges, discover improvements, generate PRDs | 15 | | /autoresearch_evals | Analyze iteration results: trends, plateaus, regressions | N/A | | /autoresearch_regression | Regression stability gate: baseline vs candidate, verdict STABLE/UNSTABLE | N/A |
| Flag | Applies To | Purpose | |---|---|---| | Iterations: N | All looping | Set iteration count | | Iterations: unlimited | All looping | Opt-in unbounded | | --evals | All looping | Mid-loop checkpoints + final summary | | --evals-interval N | All looping | Override checkpoint frequency | | --chain <targets> | All | Sequential handoff after completion | | --<subcommand> | All | Shorthand for --chain <subcommand> | | --dry-run | Orchestrator | Print derived config + planned pipeline; no execution | | --max-cycles N | Orchestrator | Hard ceiling on orchestration cycles (default 50) | | --classic | Bare /autoresearch | Force Classic metric-loop mode | | --auto | Bare /autoresearch | Force Orchestrator mode |
Activated when a plain-language goal is given without Metric:/Verify:. Classifies the goal into a Goal archetype — see references/orchestrator-routing.md for the archetype table and router decision table.
Two modes based on archetype:
Backed by scripts/orchestrate.sh (deterministic seam — all routing logic lives there). Subcommands exposed: classify, next-hop, units, plateau, screen-cmd, verdict, validate-state, screen-state-predicate.
scripts/orchestrate.sh classify "<goal>" → archetype label + mode.plan logic to produce a concrete Success predicate: exact shell command + expected output. For optimize-metric, run the full plan/wizard derivation internally.question showing: archetype, mode, concrete predicate (command + expected output), terminal choice (stop-at-verified vs proceed-to-ship). Misclassifications are caught here, not mid-run.screen-cmd; print projected cycle budget. Stop here if --dry-run.a. Assess state via cheap signals (last handoff.json, regression verdict, error count) + affected-test verify. b. scripts/orchestrate.sh next-hop orchestrator-state.json → next subcommand. c. Run subcommand (its own bounded inner loop). d. Record per-hop outcome ∈ {progressed, no-op, failed, blocked}. e. Fold hop's handoff.json into orchestrator-state.json. f. scripts/orchestrate.sh units → recompute Units remaining.
CONVERGED.scripts/orchestrate.sh plateau orchestrator-state.json → true → stop + report PLATEAU.--max-cycles N) → stop + report CEILING.blocked/failed with no alternative route → checkpoint + stop + report BLOCKED.orchestrator-state.json — orchestrator-owned, additive. Tracks: goal, archetype, predicate, terminal-choice, units_remaining history, cycle count, per-hop pipeline log with outcomes, current incumbent. Each hop's handoff.json is unchanged (single-hop bridge); the orchestrator reads it and folds it in. Two clearly-owned state objects, no overlap.
--auto to ship; deploy always requires explicit user approval.localhost/127.0.0.1/container hostname, or database name carries _test/_ci suffix. Bare substring match does not qualify. Anything else refused.screen-state-predicate and refuses on refuse.screen-cmd.orchestrator-state.json; every cycle and every resume reuses that exact string so "done" is reproducible across runs.validate-state gates orchestrator-state.json (required fields + coarse types); a malformed ledger is not trusted to route from.pending_verify; next-hop routes to a verify hop (held-out / adversarial check) before DONE or ship. The verify hop never auto-approves ship.units returns unknown (e.g. runner crash) is not counted as zero-progress; repeated unknown routes to BLOCKED.Other measured skills in the registry, with their headline benchmark lift.