Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manually trigger plan-sync to update downstream task specs after implementation drift. Use when code changes outpace specs.
.claude/skills/bilal140202-flow-next-sync/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 148% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 92% | 0% |
Manually trigger plan-sync to update downstream task specs.
CRITICAL: flowctl is BUNDLED - NOT installed globally. Define once; subsequent blocks use $FLOWCTL:
bashFLOWCTL="$HOME/.codex/scripts/flowctl" [ -x "$FLOWCTL" ] || FLOWCTL=".flow/bin/flowctl"
Non-blocking, same pattern as /flow-next:plan — one-line nag when the local setup lags the plugin:
bashSETUP_VER=$(jq -r '.setup_version // empty' .flow/meta.json 2>/dev/null) PLUGIN_JSON="${DROID_PLUGIN_ROOT:-${CLAUDE_PLUGIN_ROOT:-$HOME/.codex}}/.codex-plugin/plugin.json" PLUGIN_VER=$(jq -r '.version' "$PLUGIN_JSON" 2>/dev/null || echo "unknown") if [[ -n "$SETUP_VER" && "$PLUGIN_VER" != "unknown" && "$SETUP_VER" != "$PLUGIN_VER" ]]; then echo "Plugin updated to v${PLUGIN_VER}. Run /flow-next:setup to refresh local scripts (current: v${SETUP_VER})." >&2 fi
Continue regardless (never blocks; silent when setup was never run or versions match).
Arguments: $ARGUMENTS Format: <id> [--dry-run]
<id> - task ID fn-N-slug.M (or legacy fn-N.M, fn-N-xxx.M) or spec ID fn-N-slug (or legacy fn-N, fn-N-xxx), or a resolvable tracker handle (wor-17 / wor-17.M) that flowctl show maps to the linked spec/task (fn-52.10, R16)--dry-run - show changes without writingbashREPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
Parse $ARGUMENTS for:
ID--dry-run flag = DRY_RUN (true/false)Validate ID first (handle-recognition rule, R16):
fn-" check. Route the arg through $FLOWCTL show <ID> --json (Step 3) — flowctl's widened resolver (fn-52.10) maps a tracker key (wor-17 / wor-17.M) to its linked spec/task, so a resolvable handle is the existing spec/task, never a new id. So /flow-next:sync wor-17 resolves the linked spec.flowctl show (Step 3): "Unknown ID. Use fn-N-slug (spec) / fn-N-slug.M (task), a tracker handle (wor-17), or legacy fn-N, fn-N-xxx."Detect ID type (use the canonical id from flowctl show):
. (e.g., fn-1.2, fn-1-add-oauth.2, wor-17.2) -> task ID. (e.g., fn-1, fn-1-add-oauth, wor-17) -> spec IDbashtest -d .flow || { echo "No .flow/ found. Run flowctl init first."; exit 1; }
If .flow/ missing, output error and stop.
bash$FLOWCTL show <ID> --json
If command fails:
flowctl list to see available."flowctl specs to see available."Stop on failure.
For task ID input:
bash# Extract spec from task ID (remove .N suffix) SPEC=$(echo "<task-id>" | sed 's/\.[0-9]*$//') # Get all tasks in spec $FLOWCTL tasks --spec "$SPEC" --json
Filter to status: todo or status: blocked. Exclude the source task itself.
For spec ID input:
bash$FLOWCTL tasks --spec "<spec-id>" --json
COMPLETED_TASK_ID):status: donestatus: in_progressstatus: todo or status: blocked (these are downstream).If no downstream tasks:
No downstream tasks to sync (all done or none exist).Stop here (success, nothing to do).
Three extra context types help the agent catch drift the spec text alone can't reveal: project-glossary terms (renames where the old spec used a term whose _Avoid_ alias now appears in code), active decision constraints (current code may touch files mentioned in a decision's Consequences section), and strategic-intent drift (completed task contradicts an active STRATEGY.md track or approach).
bashGLOSSARY_JSON="$("$FLOWCTL" glossary list --json 2>/dev/null \ || echo '{"groups":[],"file_count":0,"total_terms":0}')" DECISIONS_JSON="$("$FLOWCTL" memory list --track knowledge --category decisions --json 2>/dev/null \ || echo '{"entries":[],"legacy":[],"count":0,"status":"active"}')" STRATEGY_CONTENT="$("$FLOWCTL" strategy read --json 2>/dev/null || echo '{}')"
All three calls are best-effort — empty defaults keep the agent prompt valid when flowctl returns nothing or fails.
Husk short-circuit — when ALL three of the following hold, skip the extra context entirely (pass the empty defaults; the agent's husk short-circuit at the top of Phase 3b will skip the whole section):
GLOSSARY_JSON.total_terms == 0 (glossary missing or husk)DECISIONS_JSON.count == 0 (no decision entries)STRATEGY_CONTENT.sections_filled == 0 OR STRATEGY_CONTENT == {} (no STRATEGY.md or husk — verify with flowctl strategy status --json | jq '.sections_filled // 0')When ANY of the three has signal, pass through all three (untouched) and let the agent run the matching subsection (3b.1 / 3b.2 / 3b.3) and skip the empty ones.
When GLOSSARY_JSON.total_terms == 0 but file_count > 0, every group is a husk. Husks carry no signal for drift detection — pass the JSON through untouched and let the agent skip them.
Build context and spawn via Task tool:
Sync task specs from <source> to downstream tasks.
COMPLETED_TASK_ID: <source task id - the input task, or selected source for spec mode>
FLOWCTL: $HOME/.codex/scripts/flowctl
SPEC_ID: <spec id>
DOWNSTREAM_TASK_IDS: <comma-separated list from step 4>
DRY_RUN: <true|false>
GLOSSARY_JSON: <output of `flowctl glossary list --json` from step 5>
DECISIONS_JSON: <output of `flowctl memory list --track knowledge --category decisions --json` from step 5>
STRATEGY_CONTENT: <output of `flowctl strategy read --json` from step 5>
<if DRY_RUN is true>
DRY RUN MODE: Report what would change but do NOT use Edit tool. Only analyze and report drift.
</if>Use Task tool with subagent_type: flow-next:plan-sync.
Note: COMPLETED_TASK_ID is always provided - for task-mode it's the input task, for spec-mode it's the source task selected in Step 4.
After agent returns, format output:
Normal mode:
Plan-sync: <source> -> downstream tasks
Scanned: N tasks (<list>)
<agent summary>Dry-run mode:
Plan-sync: <source> -> downstream tasks (DRY RUN)
<agent summary>
No files modified.| Case | Message | |------|---------| | No ID provided | "Usage: /flow-next:sync <id> --dry-run]" | | No .flow/ | "No .flow/ found. Run flowctl init first." | | Unknown ID (does not resolve) | "Unknown ID. Use fn-N-slug (spec) / fn-N-slug.M (task), a tracker handle (wor-17), or legacy fn-N, fn-N-xxx." | | Task not found | "Task <id> not found. Run flowctl list to see available." | | Spec not found | "Spec <id> not found. Run flowctl list to see available." | | No source (spec mode) | "No completed or in-progress tasks to sync from. Complete a task first." | | No downstream | "No downstream tasks to sync (all done or none exist)." |
planSync.enabled setting is for auto-trigger only; manual always runstodo and blocked tasks| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 9,445 | 9,708 | +3% | 1 | 1 | 0% | 1,657 | 2,991 | +81% | 0 | 0 | — |
case-02 | fail→fail | 10,639 | 7,321 | -31% | 1 | 1 | 0% | 1,941 | 2,904 | +50% | 0 | 0 | — |
case-03 | fail→fail | 8,110 | 9,015 | +11% | 1 | 1 | 0% | 1,520 | 2,879 | +89% | 0 | 0 | — |
case-04 | pass→fail | 4,504 | 7,295 | +62% | 1 | 1 | 0% | 755 | 2,828 | +275% | 0 | 0 | — |
case-05 | pass→fail | 10,404 | 5,800 | -44% | 1 | 1 | 0% | 1,931 | 2,714 | +41% | 0 | 0 | — |
case-06 | fail→fail | 4,049 | 7,078 | +75% | 1 | 1 | 0% | 182 | 2,838 | +1459% | 0 | 0 | — |
case-07 | fail→fail | 7,441 | 9,889 | +33% | 1 | 1 | 0% | 1,243 | 3,060 | +146% | 0 | 0 | — |
case-08 | fail→fail | 7,509 | 8,133 | +8% | 1 | 1 | 0% | 1,269 | 2,848 | +124% | 0 | 0 | — |
case-09 | fail→fail | 2,364 | 7,665 | +224% | 1 | 1 | 0% | 363 | 2,827 | +679% | 0 | 0 | — |
case-10 | pass→pass | 6,956 | 2,831 | -59% | 1 | 1 | 0% | 1,463 | 2,959 | +102% | 0 | 0 | — |
case-11 | fail→fail | 3,001 | 8,261 | +175% | 1 | 1 | 0% | 559 | 2,841 | +408% | 0 | 0 | — |
case-12 | fail→fail | 4,082 | 7,131 | +75% | 1 | 1 | 0% | 716 | 2,733 | +282% | 0 | 0 | — |
case-13 | fail→pass | 7,476 | 4,092 | -45% | 1 | 1 | 0% | 1,375 | 3,411 | +148% | 0 | 0 | — |
case-14 | fail→pass | 7,772 | 1,556 | -80% | 1 | 1 | 0% | 1,367 | 2,638 | +93% | 0 | 0 | — |
case-15 | fail→pass | 19,620 | 1,555 | -92% | 1 | 1 | 0% | 3,329 | 2,650 | -20% | 0 | 0 | — |
case-16 | fail→pass | 9,154 | 2,118 | -77% | 1 | 1 | 0% | 1,605 | 2,700 | +68% | 0 | 0 | — |
case-17 | fail→fail | 7,553 | 7,976 | +6% | 1 | 1 | 0% | 1,363 | 2,995 | +120% | 0 | 0 | — |
case-18 | fail→pass | 10,220 | 4,636 | -55% | 1 | 1 | 0% | 1,718 | 3,297 | +92% | 0 | 0 | — |
case-19 | fail→pass | 10,359 | 1,665 | -84% | 1 | 1 | 0% | 1,540 | 2,666 | +73% | 0 | 0 | — |
case-20 | pass→pass | 13,322 | 4,972 | -63% | 1 | 1 | 0% | 2,101 | 3,174 | +51% | 0 | 0 | — |
case-21 | fail→pass | 6,253 | 2,904 | -54% | 1 | 1 | 0% | 1,063 | 2,834 | +167% | 0 | 0 | — |
case-22 | fail→pass | 10,215 | 2,839 | -72% | 1 | 1 | 0% | 1,873 | 2,842 | +52% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 10 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.