Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Drive the canonical spec-kitty next --mission <handle> control loop for mission advancement. Load agent profiles at init, apply action-scoped doctrine context at each step boundary, and pull specific tactics/directives on demand. Triggers: "run the next step", "what should runtime do next", "advance the mission", "what is the next task", "continue the workflow", "what step comes next". Does NOT handle: setup or repair requests, purely editorial glossary or doctrine maintenance, or direct code re
.claude/skills/priivacy-ai-spec-kitty-runtime-next/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 564% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 141% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 218% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 321% | 0% |
This skill teaches agents how to advance a Spec Kitty mission through the canonical runtime control loop, including doctrine-aware context loading at each step boundary.
Use this skill when the user wants to:
The spec-kitty next command is the single entry point for agent-driven mission execution. Each call returns a deterministic decision about what action the agent should take next.
The runtime evaluates state in this order:
mission-runtime.yaml DAG)
implement and review steps, the CLI bridgemanages WP-level iteration WITHOUT advancing the runtime. The runtime only advances when ALL WPs reach terminal/handoff lanes.
first, dependency-free WPs before dependent ones
The CLI bridge (not the runtime) manages WP-level iteration:
implement or reviewplanned or in_progress lanes(done, approved, or for_review)
This means multiple calls to spec-kitty next during implementation will return different WP IDs but the same step_id (e.g., "implement") until all WPs are done.
Missions define steps as a DAG (directed acyclic graph) with dependencies:
yamlmission: key: software-dev name: Software Dev Kitty version: "2.1.0" steps: - id: discovery title: Discovery & Research depends_on: [] prompt_template: research.md - id: specify depends_on: [discovery] prompt_template: specify.md - id: plan depends_on: [specify] prompt_template: plan.md - id: tasks depends_on: [plan] - id: implement depends_on: [tasks] prompt_template: implement.md - id: review depends_on: [implement] prompt_template: review.md - id: accept depends_on: [review] prompt_template: accept.md
Every call to spec-kitty next returns exactly one decision kind:
| Kind | Meaning | Agent Action | |---|---|---| | step | Normal action available | Read prompt_file and execute | | query | Read-only current-state preview | Inspect state; do not execute a prompt or mark a result | | decision_required | Runtime needs input | Answer with --answer and --decision-id | | blocked | Guards failing, cannot proceed | Read reason + guard_failures, resolve blockers | | terminal | Mission complete | Run /spec-kitty.accept; if it passes, merge, then: mission review → author or verify retrospective (retrospect create) → surface findings (summary aggregates; synthesize reviews proposals) |
json{ "kind": "step", "agent": "claude", "mission_slug": "042-test-mission", "mission": "software-dev", "mission_state": "implementing", "action": "implement", "wp_id": "WP02", "workspace_path": ".worktrees/042-test-mission-lane-b", "prompt_file": "/tmp/spec-kitty-next-claude-042-test-mission-implement-WP02.md", "reason": null, "guard_failures": [], "progress": { "total_wps": 5, "done_wps": 1, "approved_wps": 0, "in_progress_wps": 1, "planned_wps": 3, "for_review_wps": 0 }, "run_id": "abc123", "step_id": "implement", "decision_id": null, "question": null, "options": null }
Guards block step transitions by returning failure descriptions:
| Guard | Syntax | Checks | |---|---|---| | artifact_exists | artifact_exists("spec.md") | File exists relative to mission dir | | gate_passed | gate_passed("review_gate") | Gate event in mission-events.jsonl | | all_wp_status | all_wp_status("approved_or_done") | All WPs in a specific lane or named accepted-ready set | | any_wp_status | any_wp_status("for_review") | At least one WP in lane | | input_provided | input_provided("architecture") | Input exists in runtime model | | event_count | event_count("review", 1) | Minimum event count threshold |
Guards never raise exceptions — they return false on missing context.
The runtime generates a temp file at: /tmp/spec-kitty-next-{agent}-{mission_slug}-{action}[-{wp_id}].md
Template actions (specify, plan, tasks): Mission context header + governance context + action-specific template content.
WP actions (implement, review): Full isolation-aware prompt containing:
other WPs or react to their status changes
tasks/WP##.md)Decision prompts: Question text, options, and the --answer command to run.
Runtime state is persisted between calls:
.kittify/runtime/
├── feature-runs.json # Index: {"mission-slug": {"run_id": "...", "run_dir": "..."}}
└── runs/
└── <run_id>/
└── state.json # Runtime snapshot (current step, inputs, etc.)When --mission is omitted, the runtime detects the mission via (in order):
SPECIFY_MISSION environment variable###-mission-name)NOTE: Always use --mission <slug> in multi-mission repositories.
The runtime-next loop should load doctrine context iteratively — not all at once. Each step boundary is a context loading opportunity.
At the start of a session, resolve the active agent profile. This scopes your role, boundaries, and initialization context.
Load the profile using the Python API — do NOT read YAML files directly:
pythonfrom charter.activation.doctrine_service_builder import build_activation_aware_doctrine_service service = build_activation_aware_doctrine_service(project_root) # resolve_profile()'s specializes_from lineage traversal is a repository # operation, not available on the filtered `agent_profiles` dict — reach it # through the pinned lineage/mutation accessor: repo = service.agent_profile_repository profile = repo.resolve_profile("<profile-id>") # e.g. "implementer" # Internalize identity — acknowledge this at session start print(profile.initialization_declaration) # Respect scope boundaries profile.specialization.primary_focus # What you actively do profile.specialization.avoidance_boundary # What you must NOT do profile.collaboration.handoff_to # Roles to defer to when out of scope # Load only the directives this profile references (same `service`, gated dict) for ref in profile.directive_references: directive = service.directives.get(f"DIRECTIVE_{ref.code}")
Discovery (if you don't know your profile-id):
bashspec-kitty agent profile list spec-kitty agent profile show <profile-id>
At each step boundary (when spec-kitty next returns a step decision), load governance context scoped to the current action — not the full doctrine:
bash# Load only what's relevant to this action (compact after first load) spec-kitty charter context --action implement --json
The context system uses two depth levels:
| Depth | When | Content | |---|---|---| | bootstrap (depth-2) | First load for this action | Full policy summary + reference list | | compact (depth-1) | Subsequent loads | Resolved paradigms, directives, tools only |
First-load state is tracked per action in .kittify/charter/context-state.json. This means implement and review each get their own first-load bootstrap independently.
When you need governance guidance mid-step (e.g., how to structure tests, which review criteria apply), pull the specific tactic or directive by ID rather than re-loading the full context:
pythonfrom charter.activation.doctrine_service_builder import build_activation_aware_doctrine_service service = build_activation_aware_doctrine_service(project_root) # Pull a specific tactic when it becomes relevant tactic = service.tactics.get("tdd-red-green-refactor") # Pull a specific directive directive = service.directives.get("TEST_FIRST")
The action index (actions/<action>/index.yaml) tells you which doctrine artifacts are relevant to the current step. Load the index to discover what to pull:
pythonfrom charter.offering.missions.action_index import load_action_index index = load_action_index(missions_root, "software-dev", "implement") # index.directives → ["TEST_FIRST", ...] # index.tactics → ["tdd-red-green-refactor", ...] # index.procedures → [...]
Do NOT load all doctrine into context at session start. This wastes tokens and dilutes relevance. Instead:
charter context --action <action>.Before invoking the runtime, gather the current state.
Commands:
bash# Check WP status for a mission spec-kitty agent tasks status --mission <mission-slug> # Check current context for an action spec-kitty agent context resolve --action implement --mission <mission-slug> --json
What to look for:
bash# Run the next step spec-kitty next --agent <agent> --mission <mission-slug> --json # After completing a step successfully spec-kitty next --agent <agent> --mission <mission-slug> --result success --json # After a step failed spec-kitty next --agent <agent> --mission <mission-slug> --result failed --json # After a step was blocked spec-kitty next --agent <agent> --mission <mission-slug> --result blocked --json
> Note: --mission is the sole selector. The legacy --feature alias has > been removed (#1060); passing --feature now exits with No such option.
The --result flag tells the runtime the outcome of the previous step. If omitted, spec-kitty next returns current state without advancing (query mode). It does not report success for the previous step.
See references/runtime-result-taxonomy.md for the complete taxonomy.
| Kind | Next Action | |------|-------------| | step | Read and execute prompt_file (always non-empty and resolvable on disk) | | decision_required | Answer with --answer and --decision-id | | blocked | Read reason + guard_failures, resolve blockers | | terminal | Run /spec-kitty.accept for final validation, then merge if acceptance passes |
Always check guard_failures — this field may appear on any decision kind, not just blocked.
The kind="step" prompt-file contract is a hard runtime invariant (C1/C2). A kind="step" envelope MUST carry a prompt_file (or its consumer-side prompt_path alias) that is non-null, non-empty, and resolves on disk. If the runtime cannot produce an actionable step (no composed action, guard failure, blocked dependency, prompt build error, etc.), it returns kind="blocked" with a non-empty reason (and optional machine-readable code such as no_prompt_template). There is no third state: an agent loop should never observe a kind="step" decision with prompt_file == null.
Always check progress for completion. If progress.done_wps equals progress.total_wps but kind is not terminal, the mission is actually complete (known issue #335). The runtime may not detect completion when no prior run state exists. Treat this as terminal and run /spec-kitty.accept; if acceptance passes, run /spec-kitty.merge, then run mission review and the retrospective workflow.
When the runtime needs input:
bash# The decision includes question, options, and decision_id # Answer using: spec-kitty next --agent <agent> --mission <mission-slug> \ --result success --answer "<choice>" --decision-id "<decision_id>" --json
If the agent cannot determine the answer, escalate to the user with the question and options.
See references/blocked-state-recovery.md for detailed recovery patterns.
Quick diagnostic:
bash# Check WP status and dependency graph spec-kitty agent tasks status --mission <mission-slug> # Check specific WP dependencies spec-kitty agent tasks list-dependents WP## --mission <mission-slug>
Common blockers:
| Blocker | Recovery | |---|---| | Missing artifacts (spec.md, plan.md) | Run the planning workflow first | | Upstream WP not done | Implement or review the upstream WP | | Review feedback not addressed | Re-implement, address feedback, move to for_review | | Stale agent (WP in doing, no activity) | Move WP to planned with --force | | Circular dependencies | Break cycle in WP frontmatter, re-run finalize-tasks |
The complete agent loop pattern:
bash# 1. Start the loop DECISION=$(spec-kitty next --agent claude --mission 042-mission --json) KIND=$(echo "$DECISION" | jq -r '.kind') # 2. Loop until terminal or unresolvable block while [ "$KIND" = "step" ] || [ "$KIND" = "decision_required" ]; do # Workaround #335: check progress for completion even if kind != terminal DONE=$(echo "$DECISION" | jq -r '.progress.done_wps // 0') TOTAL=$(echo "$DECISION" | jq -r '.progress.total_wps // 0') if [ "$TOTAL" -gt 0 ] && [ "$DONE" -eq "$TOTAL" ]; then break # Mission is actually complete fi if [ "$KIND" = "step" ]; then PROMPT=$(echo "$DECISION" | jq -r '.prompt_file') # Contract (C1/C2, post-#336 fix): kind=step always carries a # non-empty prompt_file resolvable on disk. If a prompt cannot be # resolved, the runtime emits kind=blocked with a populated reason. # Read and execute the prompt... RESULT="success" # or "failed" or "blocked" elif [ "$KIND" = "decision_required" ]; then # Answer the question... RESULT="success" fi DECISION=$(spec-kitty next --agent claude --mission 042-mission --result "$RESULT" --json) KIND=$(echo "$DECISION" | jq -r '.kind') done # 3. Handle terminal state — canonical post-merge sequence if [ "$KIND" = "terminal" ] || [ "$DONE" -eq "$TOTAL" ]; then # Run /spec-kitty.accept. # If acceptance passes, run /spec-kitty.merge. # After merge, follow the canonical post-merge sequence: # a. Mission review: /spec-kitty-mission-review # b. Author or verify retrospective: # spec-kitty retrospect create --mission 042-mission # if record absent # OR verify: cat .kittify/missions/<mission_id>/retrospective.yaml # c. Surface findings: # spec-kitty retrospect summary # read-only aggregation # spec-kitty agent retrospect synthesize --mission 042-mission # dry-run by default; --apply to mutate # Note: summary aggregates; synthesize applies proposals — neither authors records. fi
The loop continues until:
terminal — mission complete, exit loopquery — read-only preview, no state mutationblocked — cannot proceed without external resolutiondecision_required — only if the agent cannot answer (escalate to user)spec-kitty next rather than manually sequencing phases--mission in multi-mission repositoriesprompt_file — it contains the full context the agent needsguard_failures on every decision, not just blocked onesdownstream work
#335 — Completed missions return step instead of terminal. When spec-kitty next is called on a mission with all WPs done but no prior runtime run state, it creates a new run starting at discovery instead of recognizing the mission is complete. Workaround: Check progress.done_wps == progress.total_wps as a secondary completion signal.
#336 — fixed. prompt_file is always non-empty and resolvable on disk on kind: step decisions. When no prompt is available, the runtime now emits a structured kind: blocked decision with a non-empty reason (and optional machine-readable code such as no_prompt_template). Agent loops no longer need to defensively null-check prompt_file; a kind: step decision with a null prompt is a runtime bug.
Not all governed work happens inside an active mission. When a user asks for help with a task that has no active spec-kitty next loop — a code review, a quick implementation, an ad-hoc analysis — you should still invoke Spec Kitty's governance layer with standalone dispatch.
Spec Kitty never spawns a parallel LLM call. You are the host; Spec Kitty routes, assembles governance context, and records the trail.
If the user says anything like "use spec kitty to ...", "hey spec kitty ...", or "spec kitty <anything>", run standalone dispatch unless they are clearly asking for a full mission.
bash spec-kitty dispatch "implement the login handler" --json spec-kitty dispatch "review WP05" --profile reviewer --json Response includes invocation_id, governance_context_text, and governance_context_available.
Read governance_context_text and treat it as binding governance context for your task. Follow any directives and constraints it contains. If governance_context_available is false, note this to the user but proceed with the task.
Do the work. Generate the code, analysis, or plan.
dispatch leaves it open):bash spec-kitty profile-invocation complete \ --invocation-id <invocation_id> \ --outcome <done|failed|abandoned> Use the real outcome: done for completed work, failed for work that did not succeed, abandoned for dropped work. Never leave an Op open deliberately — spec-kitty doctor ops reports and sweeps orphans.
Every standalone invocation writes a Tier 1 JSONL file to:
kitty-ops/<invocation_id>.jsonlViewable at any time with spec-kitty invocations list --json. No SaaS connection required.
For full CLI surface documentation, see src/charter/offering/skills/spec-kitty/SKILL.md.
references/runtime-result-taxonomy.md -- Decision kinds, output fields, and precedence rulesreferences/blocked-state-recovery.md -- 6 blocked state patterns with diagnosis and recovery| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,223 | 15,812 | +4% | 1 | 1 | 0% | 269 | 5,618 | +1988% | 0 | 0 | — |
case-02 | fail→fail | 14,626 | 15,896 | +9% | 1 | 1 | 0% | 205 | 5,647 | +2655% | 0 | 0 | — |
case-03 | fail→pass | 14,941 | 25,489 | +71% | 1 | 1 | 0% | 1,281 | 8,501 | +564% | 0 | 0 | — |
case-04 | fail→pass | 19,165 | 9,375 | -51% | 1 | 1 | 0% | 2,529 | 6,104 | +141% | 0 | 0 | — |
case-05 | fail→pass | 18,070 | 8,709 | -52% | 1 | 1 | 0% | 2,221 | 6,066 | +173% | 0 | 0 | — |
case-06 | fail→pass | 15,698 | 8,952 | -43% | 1 | 1 | 0% | 1,917 | 6,090 | +218% | 0 | 0 | — |
case-07 | fail→pass | 23,778 | 10,167 | -57% | 1 | 1 | 0% | 1,475 | 6,204 | +321% | 0 | 0 | — |
case-08 | fail→pass | 13,547 | 10,471 | -23% | 1 | 1 | 0% | 1,568 | 6,463 | +312% | 0 | 0 | — |
case-09 | fail→pass | 15,002 | 10,582 | -29% | 1 | 1 | 0% | 1,755 | 6,453 | +268% | 0 | 0 | — |
case-10 | fail→pass | 14,683 | 10,612 | -28% | 1 | 1 | 0% | 1,723 | 6,438 | +274% | 0 | 0 | — |
case-11 | pass→pass | 15,301 | 9,040 | -41% | 1 | 1 | 0% | 1,702 | 6,053 | +256% | 0 | 0 | — |
case-12 | fail→pass | 20,158 | 9,778 | -51% | 1 | 1 | 0% | 2,432 | 6,276 | +158% | 0 | 0 | — |
case-13 | fail→pass | 12,474 | 3,025 | -76% | 1 | 1 | 0% | 1,459 | 5,911 | +305% | 0 | 0 | — |
case-14 | fail→pass | 19,183 | 46,825 | +144% | 1 | 1 | 0% | 2,387 | 6,193 | +159% | 0 | 0 | — |
case-15 | fail→pass | 13,781 | 7,667 | -44% | 1 | 1 | 0% | 1,673 | 5,816 | +248% | 0 | 0 | — |
case-16 | fail→pass | 15,865 | 10,391 | -35% | 1 | 1 | 0% | 1,778 | 6,290 | +254% | 0 | 0 | — |
case-17 | pass→pass | 14,750 | 9,688 | -34% | 1 | 1 | 0% | 1,617 | 6,290 | +289% | 0 | 0 | — |
case-18 | fail→pass | 18,654 | 11,909 | -36% | 1 | 1 | 0% | 2,401 | 6,704 | +179% | 0 | 0 | — |
case-19 | pass→pass | 18,918 | 13,625 | -28% | 1 | 1 | 0% | 2,190 | 6,865 | +213% | 0 | 0 | — |
case-20 | pass→pass | 17,020 | 16,189 | -5% | 1 | 1 | 0% | 2,196 | 7,385 | +236% | 0 | 0 | — |
case-21 | fail→pass | 36,397 | 10,083 | -72% | 1 | 1 | 0% | 1,661 | 6,349 | +282% | 0 | 0 | — |
case-22 | pass→pass | 17,855 | 16,758 | -6% | 1 | 1 | 0% | 2,421 | 7,548 | +212% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +68 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/2/2026 | +50% |
| gemini-3.6-flash | verified | 8/13/2026 | +64% |
Other measured skills in the registry, with their headline benchmark lift.