Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when challenging ideas, plans, decisions, or proposals using structured critical reasoning. Invoke to play devil's advocate, run a pre-mortem, red team, or audit evidence and assumptions.
.claude/skills/jeffallan-the-fool/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 12% | 0% |
The court jester who alone could speak truth to the king. Not naive but strategically unbound by convention, hierarchy, or politeness. Applies structured critical reasoning across 5 modes to stress-test any idea, plan, or decision.
AskUserQuestion with two-step mode selection (see below).Use AskUserQuestion to let the user choose how to challenge their idea.
Step 1 — Pick a category (4 options):
| Option | Description | |--------|-------------| | Question assumptions | Probe what's being taken for granted | | Build counter-arguments | Argue the strongest opposing position | | Find weaknesses | Anticipate how this fails or gets exploited | | You choose | Auto-recommend based on context |
Step 2 — Refine mode (only when the category maps to 2 modes):
references/mode-selection-guide.md and auto-recommend| Mode | Method | Output | |------|--------|--------| | Expose My Assumptions | Socratic questioning | Probing questions grouped by theme | | Argue the Other Side | Hegelian dialectic + steel manning | Counter-argument and synthesis proposal | | Find the Failure Modes | Pre-mortem + second-order thinking | Ranked failure narratives with mitigations | | Attack This | Red teaming | Adversary profile, attack vectors, defenses | | Test the Evidence | Falsificationism + evidence weighting | Claims audited with falsification criteria |
| Topic | Reference | Load When | |-------|-----------|-----------| | Socratic questioning | references/socratic-questioning.md | "Expose my assumptions" selected | | Dialectic and synthesis | references/dialectic-synthesis.md | "Argue the other side" selected | | Pre-mortem analysis | references/pre-mortem-analysis.md | "Find the failure modes" selected | | Red team adversarial | references/red-team-adversarial.md | "Attack this" selected | | Evidence audit | references/evidence-audit.md | "Test the evidence" selected | | Mode selection guide | references/mode-selection-guide.md | "You choose" selected or auto-recommend needed |
AskUserQuestion for mode selection — never assume which modeAskUserQuestion can provide structured optionsEach mode produces a structured deliverable. See the corresponding reference file for the full template.
| Mode | Deliverable | |------|------------| | Expose My Assumptions | Assumption inventory + probing questions by theme + suggested experiments | | Argue the Other Side | Steelmanned thesis + antithesis argued + synthesis proposed + confidence rating | | Find the Failure Modes | Ranked failure narratives + early warning signs + mitigations + inversion check | | Attack This | Adversary profiles + ranked attack vectors + perverse incentives + defenses | | Test the Evidence | Claims extracted + falsification criteria + evidence grades + competing explanations |
After any mode, the final output must include:
Socratic method, Hegelian dialectic, steel manning, pre-mortem analysis, red teaming, falsificationism, abductive reasoning, second-order thinking, cognitive biases, inversion technique
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,663 | 7,385 | -64% | 1 | 1 | 0% | 3,328 | 2,673 | -20% | 0 | 0 | — |
case-02 | fail→pass | 21,491 | 7,015 | -67% | 1 | 1 | 0% | 3,188 | 2,522 | -21% | 0 | 0 | — |
case-03 | fail→pass | 20,396 | 5,023 | -75% | 1 | 1 | 0% | 3,042 | 2,201 | -28% | 0 | 0 | — |
case-04 | pass→fail | 18,831 | 21,129 | +12% | 1 | 1 | 0% | 3,501 | 4,945 | +41% | 0 | 0 | — |
case-05 | pass→pass | 15,676 | 22,003 | +40% | 1 | 1 | 0% | 3,878 | 6,272 | +62% | 0 | 0 | — |
case-06 | pass→fail | 14,302 | 17,853 | +25% | 1 | 1 | 0% | 2,731 | 4,399 | +61% | 0 | 0 | — |
case-07 | pass→pass | 18,102 | 14,154 | -22% | 1 | 1 | 0% | 2,796 | 3,546 | +27% | 0 | 0 | — |
case-14 | fail→pass | 15,972 | 16,966 | +6% | 1 | 1 | 0% | 2,928 | 3,845 | +31% | 0 | 0 | — |
case-08 | fail→fail | 18,199 | 5,118 | -72% | 1 | 1 | 0% | 2,853 | 2,209 | -23% | 0 | 0 | — |
case-09 | fail→fail | 11,831 | 4,492 | -62% | 1 | 1 | 0% | 1,871 | 1,995 | +7% | 0 | 0 | — |
case-10 | pass→fail | 4,924 | 4,880 | -1% | 1 | 1 | 0% | 778 | 1,987 | +155% | 0 | 0 | — |
case-11 | fail→fail | 12,955 | 7,676 | -41% | 1 | 1 | 0% | 2,155 | 2,465 | +14% | 0 | 0 | — |
case-12 | fail→fail | 10,782 | 9,531 | -12% | 1 | 1 | 0% | 1,822 | 2,749 | +51% | 0 | 0 | — |
case-13 | fail→fail | 13,905 | 12,255 | -12% | 1 | 1 | 0% | 2,322 | 3,093 | +33% | 0 | 0 | — |
case-15 | pass→pass | 15,314 | 13,097 | -14% | 1 | 1 | 0% | 2,373 | 3,365 | +42% | 0 | 0 | — |
case-16 | fail→pass | 16,843 | 14,701 | -13% | 1 | 1 | 0% | 2,884 | 3,559 | +23% | 0 | 0 | — |
case-17 | pass→pass | 3,871 | 7,126 | +84% | 1 | 1 | 0% | 623 | 2,516 | +304% | 0 | 0 | — |
case-18 | fail→pass | 14,834 | 7,659 | -48% | 1 | 1 | 0% | 2,362 | 2,640 | +12% | 0 | 0 | — |
case-19 | fail→fail | 23,656 | 10,551 | -55% | 1 | 1 | 0% | 3,596 | 2,974 | -17% | 0 | 0 | — |
case-20 | fail→pass | 11,272 | 3,988 | -65% | 1 | 1 | 0% | 1,938 | 2,017 | +4% | 0 | 0 | — |
case-21 | fail→pass | 9,287 | 4,118 | -56% | 1 | 1 | 0% | 1,550 | 1,961 | +27% | 0 | 0 | — |
case-22 | fail→pass | 10,544 | 2,576 | -76% | 1 | 1 | 0% | 1,873 | 1,699 | -9% | 0 | 0 | — |
case-23 | fail→pass | 14,884 | 11,731 | -21% | 1 | 1 | 0% | 2,387 | 3,061 | +28% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.