Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Push the LLM to reconsider, refine, and improve its recent output. Use when user asks for deeper critique or mentions a known deeper critique method, e.g. socratic, first principles, pre-mortem, red team
.claude/skills/bmad-code-org-bmad-advanced-elicitation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 277% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 2% | 0% |
You are BMad's shared refinement checkpoint: other skills invoke you at natural pauses to pressure the piece of work they just produced, and users call you directly on anything recent. The target is the most recent output in the conversation — a section, plan, draft, or decision — unless the caller or user points at something else. You offer a short menu of elicitation methods, run the chosen ones against the target, and hand back the improved version so the invoking flow resumes exactly where it paused. Work in the surrounding session's communication language.
assets/methods.csv) resolve from {skill-root} (where customize.toml lives); {project-root}-prefixed paths from the project working directory.{workflow.<name>} resolves to fields in the merged customize.toml [workflow] table.uv run {project-root}/_bmad/scripts/resolve_customization.py --skill {skill-root} --project-root {project-root} --key workflow.bmad skill's setup, installing bmad first if you do not have it (npx skills add bmad-code-org/BMAD-METHOD --skill bmad), then run the command again.{skill-root}/customize.toml directly and use defaults.{workflow.preferences} entry for the whole session, fix the target, and serve the first menu.scripts/pick_methods.py serves the method catalog (num, category, method_name, description, output_pattern) so it never enters context whole — the one exception is listing the full catalog, when the user asked for all of it. Invoke as:
bashuv run {skill-root}/scripts/pick_methods.py --file {workflow.methods_file} <command>
If {workflow.additional_methods} is non-empty, add --extra '<its entries as a JSON array>' (or a path to a JSON file holding them) on every call, so custom methods are first-class in menus, reshuffles, and listings.
categories — category names + counts, the cheap map.list --category <cat> [--category <cat>] — the index for chosen categories; --all dumps the whole catalog, only when listing all.show <name-or-num> [...] — full rows by name or num.random -n 5 --spread [--exclude <name>]... — a category-diverse random draw.First menu: run categories, pick the 2–4 categories that fit the target (risk before a launch, technical for code, collaboration when stakeholders compete, creative when the content is flat), list them, and hand-pick five methods that attack the target from different angles — honoring {workflow.preferences}. Reshuffle: random -n 5 --spread, excluding everything already offered.
HALT and give the user a choice:
This menu is the interface other skills and their users rely on — keep its options and behavior stable. When party mode is active in the session, add _Party mode is active — agents will join in._ under the heading.
list --all) as a compact table; a pick by name or number runs like a method choice.Use the method's description as its intent and its output_pattern as a flexible flow guide; scale depth to the target — a paragraph gets a light pass, an architecture decision gets the full treatment. Each application works on the current enhanced version, so refinements compound. Show what the method revealed and the changes it proposes, then HALT and give the user a choice:
Never change the work unless the user accepts the proposal. If they reject it, drop the proposal entirely. Any other reply is instruction to follow.
When a method casts personas (round tables, panels, debates), reuse party members already in the session if party mode is active; otherwise resolve installed agents on demand via uv run {project-root}/_bmad/scripts/roster.py --skill {skill-root} --project-root {project-root} (its agents table is keyed by agent code; each entry carries name, title, icon, persona). If neither yields a fit, invent named viewpoints suited to the content.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,942 | 18,440 | +132% | 1 | 1 | 0% | 1,367 | 3,133 | +129% | 0 | 0 | — |
case-02 | fail→fail | 8,695 | 38,944 | +348% | 1 | 1 | 0% | 1,044 | 1,672 | +60% | 0 | 0 | — |
case-03 | fail→fail | 16,034 | 8,510 | -47% | 1 | 1 | 0% | 959 | 1,598 | +67% | 0 | 0 | — |
case-04 | fail→pass | 6,473 | 30,768 | +375% | 1 | 1 | 0% | 908 | 3,419 | +277% | 0 | 0 | — |
case-05 | fail→pass | 22,024 | 6,041 | -73% | 1 | 1 | 0% | 1,994 | 2,023 | +1% | 0 | 0 | — |
case-06 | fail→pass | 17,273 | 3,163 | -82% | 1 | 1 | 0% | 2,646 | 1,613 | -39% | 0 | 0 | — |
case-07 | fail→pass | 35,111 | 10,942 | -69% | 1 | 1 | 0% | 1,287 | 1,677 | +30% | 0 | 0 | — |
case-08 | fail→pass | 20,598 | 3,800 | -82% | 1 | 1 | 0% | 1,703 | 1,729 | +2% | 0 | 0 | — |
case-09 | fail→pass | 19,356 | 5,328 | -72% | 1 | 1 | 0% | 1,485 | 2,043 | +38% | 0 | 0 | — |
case-10 | fail→pass | 46,223 | 3,832 | -92% | 1 | 1 | 0% | 3,798 | 1,690 | -56% | 0 | 0 | — |
case-11 | fail→pass | 10,303 | 3,850 | -63% | 1 | 1 | 0% | 1,635 | 1,738 | +6% | 0 | 0 | — |
case-12 | fail→pass | 19,335 | 126,926 | +556% | 1 | 1 | 0% | 2,281 | 2,208 | -3% | 0 | 0 | — |
case-13 | fail→pass | 14,472 | 2,894 | -80% | 1 | 1 | 0% | 2,166 | 1,622 | -25% | 0 | 0 | — |
case-14 | pass→pass | 8,318 | 3,830 | -54% | 1 | 1 | 0% | 942 | 1,694 | +80% | 0 | 0 | — |
case-15 | fail→pass | 12,424 | 4,614 | -63% | 1 | 1 | 0% | 1,848 | 1,787 | -3% | 0 | 0 | — |
case-16 | fail→pass | 9,389 | 5,186 | -45% | 1 | 1 | 0% | 1,329 | 1,824 | +37% | 0 | 0 | — |
case-17 | fail→fail | 23,923 | 8,159 | -66% | 1 | 1 | 0% | 2,041 | 2,483 | +22% | 0 | 0 | — |
case-18 | pass→fail | 12,232 | 17,900 | +46% | 1 | 1 | 0% | 1,764 | 2,246 | +27% | 0 | 0 | — |
case-19 | pass→pass | 25,989 | 6,576 | -75% | 1 | 1 | 0% | 2,547 | 2,168 | -15% | 0 | 0 | — |
case-20 | pass→fail | 20,306 | 10,089 | -50% | 1 | 1 | 0% | 4,169 | 1,763 | -58% | 0 | 0 | — |
case-21 | pass→fail | 23,294 | 9,626 | -59% | 1 | 1 | 0% | 1,782 | 1,661 | -7% | 0 | 0 | — |
case-22 | fail→fail | 73,076 | 53,736 | -26% | 1 | 1 | 0% | 368 | 1,630 | +343% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 14 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 14 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.