Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Select and apply a complete toolkit of decision-making, problem-solving, systems-thinking, and communication models. Use when a user explicitly requests a named model or needs help framing a problem, comparing options, setting priorities, examining consequences, finding causes, estimating unknown quantities, planning toward a distant goal, stress-testing a plan, understanding a system, resolving conflict, generating solutions, giving feedback, or structuring a message. Choose the smallest useful
.claude/skills/ponomr-thinking-toolkit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 81% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 70% | 0% |
Apply structured thinking without assuming access to tools, browsing, code, memory, or a particular LLM provider. Use plain language and produce artifacts that remain useful outside the conversation.
truth or authority.
estimates, preferences, and recommendations.
only the unknowns that could change the outcome.
private hidden reasoning or produce a diary of internal deliberation.
Use the requested model when the user names it or an unambiguous alias. Read the catalog, then read only that model's card. Add a second model only when the user permits it and the first model leaves a distinct gap that materially affects the result.
Use this mode only when the user explicitly asks Thinking Toolkit or a structured thinking toolkit to choose a method but does not name one. Do not select a model merely because an ordinary request could be approached with a framework.
system, resolve conflict, give feedback, or communicate.
multiple criteria, causal ambiguity, dynamics, time pressure, or audience.
selection cues match.
has a separate role in a clear sequence.
sentences.
Capture only what matters:
Ask up to three focused questions when missing information could materially change the model, framing, or recommendation. Otherwise proceed and label reasonable assumptions.
situations. Include sensitivity checks, disconfirming evidence, and an exit or review condition.
Read the selected card before using it. Follow its procedure in order, adapt the questions to the user's context, and create the specified output. Do not reduce a model to a label or generic advice.
Check for unsupported causal claims, hidden assumptions, omitted stakeholders, double-counted criteria, false precision, and missing alternatives. Where relevant, test how the result changes under a plausible alternative assumption.
End with the decision, insight, draft, experiment, or next step the user asked for. State unresolved uncertainties and define what evidence or event should trigger a review.
The catalog is the single index for model names, aliases, selection cues, category counts, and combination recipes. Read it to resolve an explicit alias or make an authorized automatic selection, then load only the selected model card or cards.
analyze, choose, stress-test, or communicate.
workshop.
model's input.
constructing a sequence.
Match the output to the user's requested format and level of detail; no fixed set of headings is required. Include the selected model and a brief rationale when that helps the user follow the artifact. Label material assumptions and uncertainties, preserve the model's required artifact, and end with the decision, draft, insight, experiment, or next step the user requested.
/logic)The model cards above help the user choose how to think. The /logic mode does something different: it audits reasoning that already exists — a claim, an argument, or a draft — and returns a verdict on its validity.
Route here when the user explicitly invokes /logic, including with an argument, draft, or textbook logic task in any language. It has three modes:
apply a Mill's method, reconstruct an enthymeme).
Read the logic overview first — it carries the core contract, the analysis procedure, and the verdict format. Then read only the reference needed:
thought, verdict format.
modern examples.
propositional/truth-functional tests.
hypothesis strength.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 41,831 | 45,697 | +9% | 1 | 1 | 0% | 2,727 | 4,937 | +81% | 0 | 0 | — |
case-02 | fail→pass | 31,292 | 26,967 | -14% | 1 | 1 | 0% | 3,067 | 4,716 | +54% | 0 | 0 | — |
case-03 | fail→fail | 11,229 | 35,625 | +217% | 1 | 1 | 0% | 1,648 | 3,729 | +126% | 0 | 0 | — |
case-04 | pass→pass | 23,077 | 14,135 | -39% | 1 | 1 | 0% | 1,451 | 2,717 | +87% | 0 | 0 | — |
case-05 | pass→pass | 27,008 | 16,266 | -40% | 1 | 1 | 0% | 2,386 | 2,965 | +24% | 0 | 0 | — |
case-06 | pass→pass | 22,104 | 22,039 | -0% | 1 | 1 | 0% | 2,356 | 3,516 | +49% | 0 | 0 | — |
case-07 | fail→pass | 35,354 | 29,880 | -15% | 1 | 1 | 0% | 3,236 | 5,125 | +58% | 0 | 0 | — |
case-08 | fail→fail | 20,903 | 12,856 | -38% | 1 | 1 | 0% | 894 | 2,265 | +153% | 0 | 0 | — |
case-09 | fail→pass | 90,727 | 29,898 | -67% | 1 | 1 | 0% | 3,051 | 4,863 | +59% | 0 | 0 | — |
case-10 | fail→pass | 26,497 | 69,840 | +164% | 1 | 1 | 0% | 3,901 | 6,640 | +70% | 0 | 0 | — |
case-11 | fail→pass | 39,556 | 35,348 | -11% | 1 | 1 | 0% | 3,131 | 4,270 | +36% | 0 | 0 | — |
case-12 | pass→pass | 36,410 | 41,600 | +14% | 1 | 1 | 0% | 3,151 | 4,424 | +40% | 0 | 0 | — |
case-13 | pass→pass | 34,092 | 25,265 | -26% | 1 | 1 | 0% | 2,316 | 4,499 | +94% | 0 | 0 | — |
case-14 | pass→pass | 16,645 | 17,908 | +8% | 1 | 1 | 0% | 2,533 | 3,074 | +21% | 0 | 0 | — |
case-15 | pass→pass | 18,252 | 29,053 | +59% | 1 | 1 | 0% | 1,969 | 5,088 | +158% | 0 | 0 | — |
case-16 | fail→fail | 22,838 | 17,534 | -23% | 1 | 1 | 0% | 2,077 | 3,930 | +89% | 0 | 0 | — |
case-17 | pass→pass | 7,844 | 17,963 | +129% | 1 | 1 | 0% | 1,354 | 2,649 | +96% | 0 | 0 | — |
case-18 | fail→pass | 42,001 | 41,652 | -1% | 1 | 1 | 0% | 2,822 | 7,706 | +173% | 0 | 0 | — |
case-19 | fail→pass | 26,076 | 81,161 | +211% | 1 | 1 | 0% | 2,218 | 4,529 | +104% | 0 | 0 | — |
case-20 | pass→pass | 8,697 | 9,674 | +11% | 1 | 1 | 0% | 930 | 1,983 | +113% | 0 | 0 | — |
case-21 | pass→pass | 38,501 | 8,337 | -78% | 1 | 1 | 0% | 926 | 2,659 | +187% | 0 | 0 | — |
case-22 | pass→pass | 5,094 | 8,836 | +73% | 1 | 1 | 0% | 772 | 1,935 | +151% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/22/2026 | +9% |
| gemini-3.6-flash | verified | 8/3/2026 | +4% |
Other measured skills in the registry, with their headline benchmark lift.