Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optionally challenge a frozen plan with one
.claude/skills/boshu2-premortem/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 208% | 0% |
Premortem is an optional plan-challenge strategy. It asks one fresh context to identify concrete ways the resolved bead or caller intent could fail before implementation. It is not part of the required RPI sequence and does not authorize readiness.
Before any technical risk, test the plan's EVIDENCE SHAPE: for every unit of work, who verifies it, and is the verifying context distinct from the authoring context? A plan whose closure step is "the implementer runs its own tests and closes" contains no independent judgment anywhere — self-graded green is the classic false-done, and it outranks any single technical risk because it silently converts every other failure into a shipped one.
> Measured 2026-08-04, probe premortem-self-validation (gpt-5.6-luna, N=2, > directional): without this doctrine loaded the producer named the planted > self-validation flaw in 1/2 runs; with it loaded, 2/2. Ledger: > evals/skill-probes/LEDGER.md.
non-goals, evidence requirements, and declared write scope there.
reversibility, and evidence shape against cited repository facts.
Council or Dueling Idea Genies may be caller-supplied evidence, but Premortem does not require either strategy and cannot turn consensus into approval.
Actively try to construct each failure, not imagine it. For every candidate failure, attempt a concrete defeat: write the input, command sequence, or repository state that would make the plan fail, and run or cite the check that shows whether the plan survives it. A finding is reportable as concrete when it names the defeating construction and what the plan does when it lands; a failure you could not construct is reported as attempted-and-blocked with the obstacle named, which is itself evidence for the plan. The named failure mode is armchair pessimism: a list of imagined risks with no construction attempts, which reads as diligence while testing nothing. Stop condition: every reported finding is backed by a defeat attempt — constructed, or attempted with the blocking fact cited; a finding with neither is deleted, not softened.
A challenger that critiques the handed plan is a yes-man with extra steps: it anchors on the author's design and rationalizes it. Derive independently, then diff. Give one fresh context ONLY the intent source and the plan's declared ground truth — the vendor docs and stock behavior for integration work, the repo's patterns and behavior spec for extension — and never the author's design. Have it sketch its own design from that ground truth alone. The diff between that independent design and the working plan is the challenge artifact; each divergence is a finding to defend or adopt. Convergence is weak evidence the plan follows the ground truth; divergence names where it may not.
Two questions the challenger answers with an artifact, not an opinion:
exists? Artifact — the simplest version that satisfies acceptance, plus the named reason it is insufficient. No named reason means build the simple one.
counterpart in the substrate? Artifact — the native-counterpart list, one row per component the plan authors, naming the substrate feature it duplicates or the reason none exists.
These are integration- and extension-class checks. The Grain question's native-counterpart list applies only to integration-class work; do not impose it on routine feature work.
change acceptance, operate Git, close work, release, or deliver.
Return premortem-plan-review.v1 with the intent digest, author and judge context IDs, findings, evidence references, checked, and not_checked. An empty finding set means only that this optional challenge found no concrete defect; it is never a lifecycle gate.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 31,346 | 26,363 | -16% | 1 | 1 | 0% | 4,986 | 5,562 | +12% | 0 | 0 | — |
case-02 | fail→fail | 39,109 | 8,430 | -78% | 1 | 1 | 0% | 6,210 | 1,741 | -72% | 0 | 0 | — |
case-03 | fail→pass | 24,705 | 31,423 | +27% | 1 | 1 | 0% | 3,648 | 6,132 | +68% | 0 | 0 | — |
case-04 | pass→fail | 19,393 | 12,569 | -35% | 1 | 1 | 0% | 3,200 | 3,196 | -0% | 0 | 0 | — |
case-05 | pass→fail | 31,770 | 11,296 | -64% | 1 | 1 | 0% | 3,454 | 3,086 | -11% | 0 | 0 | — |
case-06 | pass→fail | 2,932 | 6,146 | +110% | 1 | 1 | 0% | 396 | 2,025 | +411% | 0 | 0 | — |
case-07 | fail→pass | 9,535 | 9,073 | -5% | 1 | 1 | 0% | 1,440 | 2,548 | +77% | 0 | 0 | — |
case-08 | fail→fail | 17,357 | 8,812 | -49% | 1 | 1 | 0% | 2,687 | 2,491 | -7% | 0 | 0 | — |
case-09 | fail→pass | 18,221 | 16,005 | -12% | 1 | 1 | 0% | 2,772 | 3,631 | +31% | 0 | 0 | — |
case-10 | fail→pass | 8,494 | 18,645 | +120% | 1 | 1 | 0% | 1,264 | 3,897 | +208% | 0 | 0 | — |
case-11 | fail→fail | 38,816 | 16,803 | -57% | 1 | 1 | 0% | 3,243 | 3,711 | +14% | 0 | 0 | — |
case-12 | pass→pass | 7,446 | 14,674 | +97% | 1 | 1 | 0% | 1,125 | 3,118 | +177% | 0 | 0 | — |
case-13 | fail→fail | 5,327 | 6,577 | +23% | 1 | 1 | 0% | 230 | 1,511 | +557% | 0 | 0 | — |
case-14 | fail→pass | 16,122 | 23,285 | +44% | 1 | 1 | 0% | 2,479 | 4,384 | +77% | 0 | 0 | — |
case-15 | fail→fail | 21,042 | 20,428 | -3% | 1 | 1 | 0% | 3,108 | 4,299 | +38% | 0 | 0 | — |
case-16 | pass→pass | 16,323 | 11,715 | -28% | 1 | 1 | 0% | 2,479 | 2,984 | +20% | 0 | 0 | — |
case-17 | fail→pass | 15,498 | 6,912 | -55% | 1 | 1 | 0% | 2,242 | 2,145 | -4% | 0 | 0 | — |
case-18 | fail→fail | 16,317 | 14,013 | -14% | 1 | 1 | 0% | 2,351 | 3,294 | +40% | 0 | 0 | — |
case-19 | fail→pass | 16,187 | 11,658 | -28% | 1 | 1 | 0% | 2,482 | 2,879 | +16% | 0 | 0 | — |
case-20 | fail→pass | 14,277 | 8,512 | -40% | 1 | 1 | 0% | 2,116 | 2,473 | +17% | 0 | 0 | — |
case-21 | pass→pass | 19,157 | 16,589 | -13% | 1 | 1 | 0% | 2,707 | 3,434 | +27% | 0 | 0 | — |
case-22 | fail→fail | 8,134 | 22,112 | +172% | 1 | 1 | 0% | 1,201 | 4,505 | +275% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 20 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.