Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Interrogate a plan, decision, or design one question at a time until it holds — use to stress-test your own thinking before committing to it
.claude/skills/hashgraph-online-skill-pressure-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -43% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 58% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 169% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 72% | 0% |
> Host: Codex CLI — This skill was designed for Claude Code and adapted for Codex. > Cross-reference commands use installed skill names in Codex rather than /octo:* slash commands. > Use the active Codex shell and subagent tools. Do not claim a provider, model, or host subagent is available until the current session exposes it. > For host tool equivalents, see skills/blocks/codex-host-adapter.md.
Interview the user relentlessly about a plan, decision, or design until you both reach a shared understanding of it. Walk each branch of the decision tree, resolving dependencies between decisions one at a time.
This is not brainstorming. skill-thought-partner opens a space up and looks for what might be there; this closes one down and looks for what is wrong with it. Use this when there is already a position on the table and the risk is that it is wrong in a way nobody has said out loud.
Adapted from the grilling skill in mattpocock/skills (MIT).
skill-decision-support.skill-thought-partner.The plan, decision, or design under test, in whatever form exists — a document, a paragraph, or just the last few turns of conversation. Nothing needs writing up first; extracting the shape is part of the job.
Ask one question at a time. Wait for the answer before asking the next. A batch of questions is a questionnaire, and it gets questionnaire answers: shallow, and shaped by whichever one the reader happened to care about. One question, answered properly, changes what the next question should be.
Carry a recommended answer with every question. "What should happen when the token expires?" is work handed back. "What should happen when the token expires? I would refresh silently and only surface an error if the refresh fails, because the alternative interrupts the user mid-task — do you agree?" is a question that can be answered in one word, and disagreed with precisely.
Look facts up; put decisions to the user. If the answer is discoverable in the filesystem, the git history, a config file, or a tool you can run, find it — asking is a tax on the user for work you could have done. Decisions are different: they are the user's, and no amount of reading the codebase produces them. Do not infer a decision from a pattern and proceed as though it were settled.
Follow dependencies, not a list. When an answer makes another question moot, drop it. When it opens two new ones, ask those before returning. The order is whatever the tree dictates.
Name disagreement when you find it. If an answer contradicts an earlier one, say which two and ask which holds. Quiet reconciliation is how a plan ends up meaning two things.
and say so when it does.
understanding. This skill produces agreement, not changes.
interrogation past that point is theatre.
to look rigorous is worse than finding none.
raise that. Do not keep grilling the details of something already dead.
later ones often depend on earlier ones.
it came from.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 11,765 | 17,045 | +45% | 1 | 1 | 0% | 1,682 | 1,595 | -5% | 0 | 0 | — |
case-02 | fail→fail | 18,439 | 11,113 | -40% | 1 | 1 | 0% | 1,909 | 1,491 | -22% | 0 | 0 | — |
case-03 | fail→fail | 12,415 | 6,850 | -45% | 1 | 1 | 0% | 1,922 | 1,554 | -19% | 0 | 0 | — |
case-04 | pass→pass | 19,884 | 33,877 | +70% | 1 | 1 | 0% | 2,567 | 4,416 | +72% | 0 | 0 | — |
case-05 | pass→fail | 22,383 | 14,490 | -35% | 1 | 1 | 0% | 2,643 | 1,506 | -43% | 0 | 0 | — |
case-06 | fail→fail | 16,105 | 5,700 | -65% | 1 | 1 | 0% | 1,331 | 1,438 | +8% | 0 | 0 | — |
case-07 | pass→pass | 11,478 | 6,781 | -41% | 1 | 1 | 0% | 1,423 | 2,335 | +64% | 0 | 0 | — |
case-08 | fail→fail | 14,115 | 7,672 | -46% | 1 | 1 | 0% | 1,384 | 1,601 | +16% | 0 | 0 | — |
case-09 | fail→fail | 20,855 | 17,197 | -18% | 1 | 1 | 0% | 2,486 | 1,465 | -41% | 0 | 0 | — |
case-10 | fail→fail | 20,657 | 15,543 | -25% | 1 | 1 | 0% | 2,607 | 1,505 | -42% | 0 | 0 | — |
case-11 | fail→fail | 7,816 | 7,549 | -3% | 1 | 1 | 0% | 1,231 | 2,259 | +84% | 0 | 0 | — |
case-12 | fail→fail | 26,070 | 18,170 | -30% | 1 | 1 | 0% | 3,441 | 1,615 | -53% | 0 | 0 | — |
case-13 | pass→pass | 12,195 | 10,140 | -17% | 1 | 1 | 0% | 2,014 | 1,904 | -5% | 0 | 0 | — |
case-14 | pass→pass | 7,245 | 10,797 | +49% | 1 | 1 | 0% | 1,097 | 2,088 | +90% | 0 | 0 | — |
case-15 | fail→pass | 6,898 | 10,458 | +52% | 1 | 1 | 0% | 961 | 1,866 | +94% | 0 | 0 | — |
case-16 | pass→fail | 11,396 | 18,579 | +63% | 1 | 1 | 0% | 981 | 1,547 | +58% | 0 | 0 | — |
case-17 | pass→pass | 12,407 | 3,737 | -70% | 1 | 1 | 0% | 871 | 1,698 | +95% | 0 | 0 | — |
case-18 | pass→pass | 11,775 | 3,802 | -68% | 1 | 1 | 0% | 1,488 | 1,644 | +10% | 0 | 0 | — |
case-19 | pass→pass | 14,730 | 9,710 | -34% | 1 | 1 | 0% | 2,319 | 2,270 | -2% | 0 | 0 | — |
case-20 | pass→pass | 10,738 | 10,274 | -4% | 1 | 1 | 0% | 1,622 | 1,732 | +7% | 0 | 0 | — |
case-21 | pass→fail | 9,231 | 20,973 | +127% | 1 | 1 | 0% | 552 | 1,486 | +169% | 0 | 0 | — |
case-22 | fail→fail | 33,280 | 19,781 | -41% | 1 | 1 | 0% | 5,178 | 1,353 | -74% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 10 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.