Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Adversarial pre-implementation pass — argue against a proposed plan before writing code, and emit a binding PROCEED / REVISE / REDESIGN verdict
.claude/skills/nudgebee-challenge/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 15% | 0% |
<!-- Single-source-of-truth skill: this file is canonical under .claude/skills/challenge/. .gemini/skills/challenge is a directory symlink to this directory. Both agents parse the YAML frontmatter above — Claude reads user-invocable + allowed-tools, Gemini reads name, and both tolerate the extra fields. If you edit this file, both agents pick up the change automatically. Do NOT copy this skill into .gemini/skills/ — use the symlink. -->
Run an adversarial pass against a proposed plan before any code is written. This catches wrong-direction work before it accumulates sunk cost.
The failure mode we are defending against: the AI writing exactly what was asked for — when what was asked for was wrong.
Takes a plan description via $ARGUMENTS. If $ARGUMENTS is empty, challenge the most recently discussed plan in the current session (or the current branch's diff if there is already code).
Use aggressively for:
CLAUDE.md → Database Migrations & RPC Actions)Skip for:
When in doubt, run it. The cost is five minutes; the cost of building the wrong thing is much higher.
Before challenging, confirm the plan is concrete enough to attack:
If any of those are missing, stop and ask the user to sharpen the plan first. An ambiguous plan cannot be meaningfully challenged — you will end up arguing with straw men.
In one paragraph, restate the proposed approach as precisely as possible. Include: the problem it solves, the approach, the affected services, and how you'd know it worked.
Produce exactly three independent reasons this plan is wrong. Not nitpicks — structural objections. Each counterargument must include:
Bar for a valid counterargument:
If you cannot find three, you either (a) don't understand the plan well enough, or (b) the plan is actually fine for a trivial change — in which case, skip this skill.
Answer these three questions plainly:
End with exactly one of these three verdicts. The verdict is binding. Implementation MUST NOT start on anything but PROCEED.
PROCEED — The objections are real but the plan is still the right call. Note which objections to actively mitigate during implementation.REVISE — The plan mostly holds but needs specific changes before implementation. List the exact required changes. Re-run /challenge on the revised plan if any change is material.REDESIGN — One or more counterarguments are fatal, or a simpler alternative dominates. Do NOT implement this plan. Return to the research/strategy phase.If the verdict is PROCEED and the change is architectural (affects shared contracts, schema, cross-service behavior, framework choice, or tooling), append one line to the root CLAUDE.md under ## AI Coding Principles → ### 4. Decisions & Lessons Learned → #### Architecture Decisions (removing the _No entries yet_ placeholder if this is the first entry). Format:
- **[YYYY-MM] {title}**: Chose {approach} over {alternative}. Why: {reason}. Counterarguments considered: {brief summary}. Reconsider if: {condition}.If the verdict is REDESIGN and the plan had been seriously considered (not just a half-formed idea), append to ## Decisions & Lessons Learned → What We've Tried and Won't Try Again:
- **[YYYY-MM] {title}**: Considered {approach} for {problem}. Rejected because: {concrete reason from counterarguments}. Current direction: {replacement}.Do NOT log day-to-day implementation details — only decisions that affect how future work should be done.
Always output in this shape so the result is parseable and reviewable:
## Adversarial Review: {plan title}
### Restated Plan
{one paragraph}
### Counterargument 1 — {short title}
- **Objection:** ...
- **Risk:** ...
- **Cost of ignoring:** ...
### Counterargument 2 — {short title}
- **Objection:** ...
- **Risk:** ...
- **Cost of ignoring:** ...
### Counterargument 3 — {short title}
- **Objection:** ...
- **Risk:** ...
- **Cost of ignoring:** ...
### Future-Self Review
- **Senior-engineer critique (6 months out):** ...
- **Optimizes for:** ...
- **Sacrifices:** ...
- **Simpler alternative:** ... (or: "none found — full approach justified because ...")
### Verdict: {PROCEED | REVISE | REDESIGN}
{For PROCEED: which objections must be actively mitigated during implementation.}
{For REVISE: the exact required changes. Re-run /challenge if material.}
{For REDESIGN: a one-line pointer to what to explore instead.}PROCEED on every run means the skill is broken. If you haven't emitted REVISE or REDESIGN in a while, check whether you're being critical enough.REVISE or REDESIGN. Stop. Surface the verdict to the user. Let them decide the next move.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | fail→pass | 12,698 | 18,099 | +43% | 1 | 1 | 0% | 2,167 | 4,677 | +116% | 0 | 0 | — |
case-17 | fail→pass | 21,868 | 19,785 | -10% | 1 | 1 | 0% | 3,489 | 5,136 | +47% | 0 | 0 | — |
case-18 | fail→fail | 22,897 | 18,070 | -21% | 1 | 1 | 0% | 3,796 | 4,576 | +21% | 0 | 0 | — |
case-01 | fail→pass | 26,671 | 25,016 | -6% | 1 | 1 | 0% | 4,266 | 3,973 | -7% | 0 | 0 | — |
case-02 | fail→pass | 34,819 | 19,307 | -45% | 1 | 1 | 0% | 2,879 | 4,447 | +54% | 0 | 0 | — |
case-03 | fail→fail | 25,010 | 7,040 | -72% | 1 | 1 | 0% | 3,718 | 2,127 | -43% | 0 | 0 | — |
case-04 | fail→pass | 16,108 | 9,461 | -41% | 1 | 1 | 0% | 2,476 | 2,850 | +15% | 0 | 0 | — |
case-05 | fail→fail | 70,776 | 37,737 | -47% | 1 | 1 | 0% | 2,565 | 8,024 | +213% | 0 | 0 | — |
case-06 | pass→pass | 9,599 | 3,239 | -66% | 1 | 1 | 0% | 1,562 | 2,299 | +47% | 0 | 0 | — |
case-07 | fail→fail | 9,581 | 3,287 | -66% | 1 | 1 | 0% | 1,407 | 2,177 | +55% | 0 | 0 | — |
case-08 | pass→pass | 8,739 | 2,682 | -69% | 1 | 1 | 0% | 1,352 | 2,170 | +61% | 0 | 0 | — |
case-09 | pass→pass | 18,310 | 22,475 | +23% | 1 | 1 | 0% | 3,212 | 4,478 | +39% | 0 | 0 | — |
case-10 | fail→pass | 19,041 | 24,523 | +29% | 1 | 1 | 0% | 3,015 | 5,969 | +98% | 0 | 0 | — |
case-12 | fail→fail | 18,640 | 8,684 | -53% | 1 | 1 | 0% | 2,615 | 2,280 | -13% | 0 | 0 | — |
case-13 | fail→pass | 12,159 | 51,479 | +323% | 1 | 1 | 0% | 1,896 | 5,639 | +197% | 0 | 0 | — |
case-14 | fail→pass | 21,189 | 17,776 | -16% | 1 | 1 | 0% | 3,245 | 4,658 | +44% | 0 | 0 | — |
case-15 | fail→fail | 19,014 | 7,928 | -58% | 1 | 1 | 0% | 2,939 | 3,205 | +9% | 0 | 0 | — |
case-16 | fail→fail | 13,843 | 16,868 | +22% | 1 | 1 | 0% | 2,079 | 4,689 | +126% | 0 | 0 | — |
case-19 | fail→fail | 21,953 | 5,396 | -75% | 1 | 1 | 0% | 3,339 | 2,171 | -35% | 0 | 0 | — |
case-20 | fail→pass | 36,036 | 12,093 | -66% | 1 | 1 | 0% | 5,662 | 3,498 | -38% | 0 | 0 | — |
case-21 | pass→pass | 21,325 | 12,150 | -43% | 1 | 1 | 0% | 3,220 | 3,589 | +11% | 0 | 0 | — |
case-22 | fail→pass | 31,544 | 21,535 | -32% | 1 | 1 | 0% | 4,712 | 4,920 | +4% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/22/2026 | +64% |
Other measured skills in the registry, with their headline benchmark lift.