Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Garden a repo's plan portfolio from accumulated evidence — detect closures (a queue item whose phases now `dos verify` as shipped is done), track cooldown state, and surface the 0-2 items the operator must actually decide via the `dos decisions` queue. Read-only on code/data; writes only its queue + cooldown state. Driven by `dos` verbs + the workspace's `dos.toml`; names no host path or convention. Use after a burst of dispatches, when the backlog looks drained, or when recurring findings start
.claude/skills/bilal140202-dos-replan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 400% | 0% |
> The planning refresh that keeps the loop honest. Its domain-free core is > small: a queue item whose phases now dos verify as shipped is closed; the > operator should see only the 0-2 items that truly need a human. The heavy > host-specific gardening passes (anchor reconciliation, soak-state drift, the > postmortem evidence stream) are host driver hooks, scoped OUT of the generic > baseline. Closure detection rides the truth syscall; the operator surface is > the dos decisions queue.
The shape: read the queue → close what verifies as shipped → rank what's left → surface only what needs a human. Closure is a kernel call (dos verify); the operator inbox is the kernel's decisions projection.
--force (optional) — run even if the no-op skip gate would fire (no newevidence since the last sweep).
bashdos doctor --workspace . --json
Read paths (where the queue/cooldown state live) and stamp (the active ship grammar dos verify applies). Read the active trunk from config if you will gate a release later (see /dos-replan-loop).
If there is no new evidence since the last sweep (no new commits, no new findings) and --force was not given, print one line and exit cheap — writing nothing. This is the gate that keeps a recurring loop from doing 0-work sweeps; an unproductive replan (this skip, or a 0/0/0 sweep) must NOT arm the loop's drained-twice trigger (the kernel loop decision enforces that — report the productivity honestly so /dos-dispatch-loop reads it right).
For each open item in the queue, ask the kernel whether its phases have shipped:
bashdos verify --workspace . <PLAN> <PHASE> --json
If every phase of a queue item now reports shipped: true, the item is closed — move it from the open queue to the closed section. Never trust the plan doc's own stamp; the truth syscall is the source. This is the auto-close pass that keeps the queue from carrying already-shipped work as if it were pending.
Count the closures — this is part of the productivity signal Step 5 reports.
Bump the cooldown timestamp in the configured state file so /dos-next-up knows when this sweep last ran (the cooldown banner). Generic, no host specifics.
Order the remaining open items by a domain-free signal (recency, how many phases remain). Do not impose a host's bespoke ranking — that's a driver hook.
The ruthless filter: of everything you swept, at most 0-2 items should reach the operator — the ones no mechanism can resolve (a real decision-needed). The kernel's operator inbox is the dos decisions queue; read it to see what is already pending so you don't duplicate a row:
bashdos decisions --workspace . --json
The dos decisions projection is the generic operator inbox: a HUMAN-resolvable row is one a mechanism (an ORACLE/JUDGE) cannot close.
Honesty about the write path (a named open seam). dos decisions is read-only today — there is no generic dos verb to append a decision-needed row (the queue's write path, home.append_decision, is currently reached only by dos arbitrate --force's override capture). So the generic baseline surfaces the 0-2 items in its operator summary (Step 6) and reads the existing queue here; emitting a new decision row into the queue is a host/driver capability (the named open seam — see docs/74-friction-log.md). Do not imply dos decisions writes; it lists and drills in.
A sweep surfaces more than closures: a bug in another lane, a missing test, a doc that drifted from the code. The 0-2 operator slots are for decisions; a concrete, public-subject finding with a checkable done-condition is work, and its durable home is the workspace's public issue tracker (on a GitHub-hosted repo, the gh CLI) — not the inbox, not a memory file, not silence. The filing discipline is /dos-dispatch's "Out-of-scope findings" section: dedupe first (gh issue list --search "<keywords>"); give the body a checkable done-condition, a lane guess, and where you found it; leak-check the drafted body before posting (issue text is public output no tracked-file gate scans — no machine-absolute paths, hostnames, or personal identifiers; if the workspace ships a publication leak-scanner, pipe the draft through it and treat a hit as a refusal). An issue closes only via Fixes #N in the body of the commit that resolves it — never gh issue close off your own narration.
Print a terse summary: how many items closed (Step 2), how many remain, and the 0-2 that need a decision (with a pointer to dos decisions). Report the productivity verdict honestly (PRODUCTIVE iff it closed/refilled/gardened anything) — /dos-dispatch-loop reads this to decide drained-twice.
postmortem evidence stream, gitignore drift — those are host-specific evidence surfaces, scoped OUT of the generic baseline (a driver hook, or out of scope). The generic sweep does closure detection + the operator surface, and logs what it is skipping.
dos decisions queue — thegeneric operator-inbox surface — not a host-specific pending-decisions file.
> The stale-stamp catch. A queue item points at a plan whose phase the doc > still narrates as in-flight — but the grep rung sees a commit subject carrying > the phase token, so dos verify flips it SHIPPED. Read the rung, not the > bare verdict, then reconcile the queue.
Step 0 — discover the layout (the on-ramp every sweep runs):
bash$ dos doctor --workspace . --json ... "stamp": {"style": "grep", "grammar": "generic (any/no dir prefix)"} ... ... "exit_codes": {"gate": {"LIVE": 0, "DRAIN": 3, "STALE-STAMP": 4, "BLOCKED": 5, "RACE": 6}} ...
Step 2 — closure detection. Ask the truth syscall, never the plan doc's own stamp:
bash$ dos verify --workspace . docs/82_liveness-oracle-plan liveness --json {"phase":"liveness","plan":"docs/82_liveness-oracle-plan","rung":"direct","sha":"80d4f30","shipped":true,"source":"grep-subject","summary":"80d4f30 liveness: exclude the BIRTH acquire from the ADVANCING event count"}
Verdict carried by the source: grep-subject rung — a commit SUBJECT containing the phase token flips this to SHIPPED even if little was built. The doc lagging the git fact IS the stale stamp; close the item from this rung, not from the doc.
Contrast — an item still genuinely in flight returns the none rung:
bash$ dos verify --workspace . docs/99_runtime-validation-and-the-actuation-boundary halt --json {"phase":"halt","plan":"docs/99_runtime-validation-and-the-actuation-boundary","shipped":false,"source":"none"}
Step 6 reconcile — feed the dispositions to the gate; a stale stamp returns exit 4:
bash$ dos gate ./dispositions.json ; echo "exit=$?" exit=4
gate exit 4 = STALE-STAMP — the plan doc lags the verified git fact; the replan's reconciliation is to close the over-claimed item against the grep-subject rung and drop the doc's self-narrated status. (Exit 0 = LIVE, 3 = DRAIN, 5 = BLOCKED.)
dos verify only.everything a mechanism can resolve stays out of it.
drained-twice stop in the loop. Report honestly.
inbox is for decisions; file the finding as an issue with a done-condition.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,249 | 5,196 | -1% | 1 | 1 | 0% | 282 | 2,563 | +809% | 0 | 0 | — |
case-11 | fail→pass | 10,405 | 3,359 | -68% | 1 | 1 | 0% | 1,834 | 2,701 | +47% | 0 | 0 | — |
case-02 | fail→fail | 3,707 | 6,482 | +75% | 1 | 1 | 0% | 247 | 2,520 | +920% | 0 | 0 | — |
case-03 | fail→fail | 5,965 | 13,595 | +128% | 1 | 1 | 0% | 308 | 4,846 | +1473% | 0 | 0 | — |
case-04 | pass→pass | 7,548 | 3,408 | -55% | 1 | 1 | 0% | 1,307 | 2,691 | +106% | 0 | 0 | — |
case-05 | fail→pass | 9,120 | 3,682 | -60% | 1 | 1 | 0% | 1,446 | 2,851 | +97% | 0 | 0 | — |
case-06 | fail→fail | 6,512 | 2,883 | -56% | 1 | 1 | 0% | 1,018 | 2,741 | +169% | 0 | 0 | — |
case-07 | fail→pass | 8,529 | 5,047 | -41% | 1 | 1 | 0% | 1,376 | 3,141 | +128% | 0 | 0 | — |
case-08 | fail→pass | 16,909 | 4,460 | -74% | 1 | 1 | 0% | 2,561 | 2,916 | +14% | 0 | 0 | — |
case-09 | fail→fail | 5,295 | 5,874 | +11% | 1 | 1 | 0% | 188 | 2,316 | +1132% | 0 | 0 | — |
case-10 | fail→pass | 4,103 | 3,622 | -12% | 1 | 1 | 0% | 545 | 2,724 | +400% | 0 | 0 | — |
case-12 | fail→fail | 12,518 | 3,527 | -72% | 1 | 1 | 0% | 2,095 | 2,644 | +26% | 0 | 0 | — |
case-13 | fail→pass | 4,352 | 2,167 | -50% | 1 | 1 | 0% | 641 | 2,498 | +290% | 0 | 0 | — |
case-14 | fail→pass | 8,417 | 3,133 | -63% | 1 | 1 | 0% | 1,408 | 2,685 | +91% | 0 | 0 | — |
case-15 | fail→fail | 9,907 | 4,270 | -57% | 1 | 1 | 0% | 1,476 | 2,888 | +96% | 0 | 0 | — |
case-16 | pass→pass | 7,139 | 5,072 | -29% | 1 | 1 | 0% | 1,044 | 3,072 | +194% | 0 | 0 | — |
case-17 | pass→pass | 8,096 | 4,194 | -48% | 1 | 1 | 0% | 1,270 | 2,748 | +116% | 0 | 0 | — |
case-18 | fail→pass | 13,113 | 5,629 | -57% | 1 | 1 | 0% | 2,256 | 3,099 | +37% | 0 | 0 | — |
case-19 | fail→pass | 18,850 | 3,454 | -82% | 1 | 1 | 0% | 1,486 | 2,686 | +81% | 0 | 0 | — |
case-20 | fail→fail | 8,601 | 2,314 | -73% | 1 | 1 | 0% | 1,403 | 2,521 | +80% | 0 | 0 | — |
case-21 | fail→pass | 10,252 | 4,349 | -58% | 1 | 1 | 0% | 1,715 | 3,001 | +75% | 0 | 0 | — |
case-22 | fail→fail | 25,367 | 2,755 | -89% | 1 | 1 | 0% | 1,543 | 2,637 | +71% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.