Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Render a cognitive-aid PR body from flow-next state and open via gh. Triggers on /flow-next:make-pr with optional spec id and flags (--draft, --ready, --no-mermaid, --base <ref>, --memory, --dry-run). Auto-detects spec from current branch when no id given. NOT Ralph-blocked — autonomous loops can surface a draft PR for human review.
.claude/skills/bilal140202-flow-next-make-pr/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 162% | 0% |
A reviewable PR body is itself an artefact: it lets a human decide where to focus before skimming the diff. flow-next already collects every input that body needs — the spec with R-IDs, per-task done summaries and evidence commits, decisions / bug / architecture-patterns memory entries, glossary changes, strategy alignment, deferred review findings, the git diff itself. This skill stitches those into a structured body, optionally adds mermaid diagrams for module-boundary changes, and pushes via gh pr create.
The host agent (Claude Code / Codex / Droid) reads the structured payload from flowctl spec export-cognitive-aid and synthesizes the body directly. Every claim in the body must trace to a structured field in the export payload — never fabricate file paths, SHAs, R-ID attributions, or "why" reasoning. Unknown attribution is honest ("uncovered" / "unclear") rather than invented. The host is competent at "what looks important here?" given the rich input; no second-model review pass is needed (the structured payload does the heavy lifting).
flowctl provides only thin plumbing: flowctl spec export-cognitive-aid <spec-id> --base <ref> --json aggregates the inputs into a single JSON payload (Task 1 of this spec). The skill renders the body, then pushes and creates the PR directly — no confirm prompt (invoking make-pr is the intent; the body is deterministic; the default is a reversible draft). --dry-run prints the body without creating; --ready/--draft set draft state.
Read workflow.md for the full phase-by-phase execution. Read phases.md for the per-phase Done-when checklists. Read mermaid-rules.md before emitting any mermaid codefence — it defines reserved words, escape patterns, shape selection, and the pre-emission validation checklist.
CRITICAL: flowctl is BUNDLED — NOT installed globally. which flowctl will fail (expected). Define once; subsequent blocks (here and in workflow.md / phases.md) use $FLOWCTL:
bashFLOWCTL="${DROID_PLUGIN_ROOT:-${CLAUDE_PLUGIN_ROOT}}/scripts/flowctl"
Inline skill (no context: fork) — AskUserQuestion must stay reachable for the Phase 0 info prompts (resolve a missing base ref / undetected spec id — never a confirm gate). Subagents can't call blocking question tools (Claude Code issues #12890, #34592). There is no Phase 4 confirm prompt — make-pr creates the PR directly. (sync-codex.sh rewrites any remaining AskUserQuestion to a plain-text numbered prompt in the Codex mirror.)
Parse $ARGUMENTS as a flag list. Recognized flags: --draft, --ready, --no-mermaid, --memory, --dry-run, and --base <ref> (consumes the next token). Strip recognized tokens; the remainder (if any) is the optional spec id.
bashRAW_ARGS="$ARGUMENTS" DRAFT_FORCE="auto" # auto | draft | ready NO_MERMAID=0 WRITE_MEMORY=0 DRY_RUN=0 BASE_REF="" SPEC_ID="" # Tokenize and walk the argument list. set -- $RAW_ARGS while [[ $# -gt 0 ]]; do case "$1" in --draft) DRAFT_FORCE="draft"; shift ;; --ready) DRAFT_FORCE="ready"; shift ;; --no-mermaid) NO_MERMAID=1; shift ;; --memory) WRITE_MEMORY=1; shift ;; --dry-run) DRY_RUN=1; shift ;; --base) BASE_REF="$2"; shift 2 ;; --base=*) BASE_REF="${1#--base=}"; shift ;; --) shift; break ;; -*) echo "Unknown flag: $1" >&2; exit 2 ;; *) SPEC_ID="$1"; shift ;; esac done
| Flag | Effect | |------|--------| | --draft | Force draft PR regardless of open-items count or Ralph context. | | --ready | Force non-draft PR. Conflicts with --draft (last flag wins; surface the conflict). | | --no-mermaid | Skip Phase 3 entirely. Mermaid prose summaries are also skipped. | | --memory | After PR creation, write a knowledge/architecture-patterns/ memory entry summarizing what shipped. Idempotent — rerun adds no second entry for the same spec id. | | --dry-run | Skip Phase 4 entirely. Render body to stdout. Useful for inspection or … --dry-run \| pbcopy. | | --base <ref> | Override base-branch detection cascade. Useful when the team's default branch is develop, etc. |
Ralph mode (FLOW_RALPH=1 or REVIEW_RECEIPT_PATH set) is detected separately in workflow.md §0.0 — the skill is not Ralph-blocked. Under Ralph the skill hard-errors instead of asking the Phase 0 info prompts, forces --draft, and emits the PR URL to stdout. (The PR is created directly in both modes — the only difference is forced-draft + no Phase 0 prompts under Ralph.)
AskUserQuestion (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded). Fall back to a numbered options prompt only if the tool is unreachable. Never silently skip the question.--base and no detection match; no spec detected) — never "do you want to create it?". Not-all-tasks-done warns and proceeds (the open items make it a draft). Skip questions when context resolves cleanly.The body is synthesized from the export payload. Every claim must trace to a structured field. The skill explicitly forbids:
git diff --name-status (via the diff.files array) appear in Critical Changes / Where to look. No "I think there's also a config file" content.tasks[].evidence[].commits and git log --oneline base..HEAD only.satisfies frontmatter and commit-message R-ID references. Uncovered R-IDs get a ⚠️ flag, never a confident attribution.memory.decisions[] entries' bodies. If no decision entry exists for a change, the body says so explicitly rather than narrating a plausible-sounding rationale.reviews.deferred[] and reviews.suppressed_count. The body never editorializes severity or fabricates findings.strategy.tracks[] and the spec's ## Strategy Alignment block. The body never invents alignment claims.glossary.changes[]. New terms / renamed terms are surfaced only if the export reports them.git diff analysis (Phase 3 details in the mermaid-rules.md ref file). The skill never adds "I think module X also imports Y" edges.When data is missing, the body says so honestly (e.g. *No decision-track memory entries for this spec. Surface decisions in PR review comments if needed.*) rather than confabulating content. Honest "unclear" beats plausible "wrong".
--draft forced). Do NOT add a FLOW_RALPH/REVIEW_RECEIPT_PATH exit-2 guard at the top of the skill.AskUserQuestion before push. The escape hatch is --dry-run, not a question.--dry-run mode. Phase 4 short-circuits before any git push or gh pr create. The body lands on stdout only.gh pr view --json url 2>/dev/null returns rc=0 for CLOSED and MERGED PRs as readily as OPEN. Filter .state == "OPEN" via jq (validated empirically during fn-42 spike). Closed/merged PRs on a reused branch must NOT trigger refusal.git push workflows when gh is missing. When gh isn't installed or authenticated, surface the install / gh auth login instructions and exit. Don't try to fall back to half-baked PR creation.--memory. Default off. The user opts in for structurally-significant specs — every-PR memory inflation is the failure mode this gate prevents.gh pr merge. Out of scope. The skill creates and exits; merge is a human decision.Same pattern as /flow-next:plan and /flow-next:audit — non-blocking notice when .flow/meta.json setup_version lags the plugin version:
bashif [[ -f .flow/meta.json ]]; then SETUP_VER=$(jq -r '.setup_version // empty' .flow/meta.json 2>/dev/null) PLUGIN_JSON="${DROID_PLUGIN_ROOT:-${CLAUDE_PLUGIN_ROOT}}/.claude-plugin/plugin.json" PLUGIN_VER=$(jq -r '.version' "$PLUGIN_JSON" 2>/dev/null || echo "unknown") if [[ -n "$SETUP_VER" && "$PLUGIN_VER" != "unknown" && "$SETUP_VER" != "$PLUGIN_VER" ]]; then echo "Plugin updated to v${PLUGIN_VER}. Run /flow-next:setup to refresh local scripts (current: v${SETUP_VER})." >&2 fi fi
Execute the phases in workflow.md in order:
gh installed + authenticated; resolve spec id (arg or branch-match); base-branch detection cascade; branch validity (HEAD ahead of base); all tasks done (warn + proceed as draft if not — no prompt; Ralph exits 2); existing-PR refusal filtered on .state == "OPEN". Detects Ralph environment for downstream phases.flowctl spec export-cognitive-aid <spec-id> --base <ref> --json; parse the structured payload (spec / tasks / memory / glossary / strategy / diff / reviews).mermaid-rules.md §6 checklist (reserved words, escape patterns, no emoji / MathJax, no inheritance cycles) before emitting. Skipped under --no-mermaid or when no triggers fire / a skip rule applies (pure-additive single-module diff <50 LOC, flat-layout repo).git push -u origin HEAD, then gh pr create --title --body. Draft when OPEN_ITEMS_COUNT > 0 OR Ralph OR --draft; ready when --ready. --dry-run short-circuits before push.Generated by /flow-next:make-pr from <spec-id> against <base>); optionally write knowledge/architecture-patterns/ memory entry under --memory.Phase 0 is implemented in this task (fn-42.2). Phases 1-5 land in fn-42.3 → fn-42.6 (body sections, mermaid, push). The skill scaffold here owns the structure; per-phase content is filled in dependent tasks.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→pass | 10,813 | 3,092 | -71% | 1 | 1 | 0% | 2,128 | 3,694 | +74% | 0 | 0 | — |
case-01 | fail→fail | 13,333 | 6,001 | -55% | 1 | 1 | 0% | 2,443 | 3,430 | +40% | 0 | 0 | — |
case-02 | fail→fail | 9,762 | 9,267 | -5% | 1 | 1 | 0% | 1,612 | 3,828 | +137% | 0 | 0 | — |
case-03 | fail→fail | 3,942 | 7,551 | +92% | 1 | 1 | 0% | 297 | 3,515 | +1084% | 0 | 0 | — |
case-04 | fail→fail | 2,985 | 10,069 | +237% | 1 | 1 | 0% | 487 | 5,039 | +935% | 0 | 0 | — |
case-05 | fail→fail | 10,618 | 5,194 | -51% | 1 | 1 | 0% | 1,996 | 3,277 | +64% | 0 | 0 | — |
case-06 | fail→fail | 2,550 | 6,380 | +150% | 1 | 1 | 0% | 304 | 3,591 | +1081% | 0 | 0 | — |
case-07 | fail→pass | 10,242 | 2,000 | -80% | 1 | 1 | 0% | 1,944 | 3,474 | +79% | 0 | 0 | — |
case-09 | fail→pass | 11,771 | 2,744 | -77% | 1 | 1 | 0% | 2,428 | 3,723 | +53% | 0 | 0 | — |
case-10 | pass→pass | 9,492 | 4,908 | -48% | 1 | 1 | 0% | 1,554 | 4,084 | +163% | 0 | 0 | — |
case-11 | pass→pass | 3,236 | 1,888 | -42% | 1 | 1 | 0% | 499 | 3,435 | +588% | 0 | 0 | — |
case-12 | pass→pass | 6,676 | 4,020 | -40% | 1 | 1 | 0% | 1,236 | 3,830 | +210% | 0 | 0 | — |
case-13 | fail→pass | 9,135 | 3,953 | -57% | 1 | 1 | 0% | 1,612 | 3,796 | +135% | 0 | 0 | — |
case-14 | pass→pass | 11,303 | 2,688 | -76% | 1 | 1 | 0% | 1,703 | 3,639 | +114% | 0 | 0 | — |
case-15 | fail→pass | 8,415 | 1,980 | -76% | 1 | 1 | 0% | 1,324 | 3,463 | +162% | 0 | 0 | — |
case-16 | fail→pass | 8,274 | 3,582 | -57% | 1 | 1 | 0% | 1,445 | 3,807 | +163% | 0 | 0 | — |
case-17 | pass→pass | 4,272 | 2,312 | -46% | 1 | 1 | 0% | 689 | 3,521 | +411% | 0 | 0 | — |
case-18 | pass→pass | 7,286 | 2,474 | -66% | 1 | 1 | 0% | 1,450 | 3,548 | +145% | 0 | 0 | — |
case-19 | pass→pass | 6,407 | 3,269 | -49% | 1 | 1 | 0% | 1,045 | 3,622 | +247% | 0 | 0 | — |
case-20 | fail→fail | 8,886 | 1,915 | -78% | 1 | 1 | 0% | 1,367 | 3,404 | +149% | 0 | 0 | — |
case-21 | fail→pass | 10,159 | 2,468 | -76% | 1 | 1 | 0% | 1,914 | 3,699 | +93% | 0 | 0 | — |
case-22 | pass→pass | 11,353 | 3,770 | -67% | 1 | 1 | 0% | 1,989 | 3,860 | +94% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.