Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing results and discussion, and when a task or a reviewer asks why an effect happens rather than whether it does. Covers the difference between reporting an effect and accounting for it, and what a mechanism claim needs behind it.
.claude/skills/tangxiangru-answer-the-why-not-only-the-what/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 9% | 0% |
Plenty of the questions a study is judged on are not "what is the number" but "why is it that number": why two methods that should differ give the same gain, why an effect is larger in one regime, why the improvement vanishes past a threshold, what the failure cases have in common.
A report can measure such an effect precisely and score nothing on it, because it never says what the effect is evidence of. Numbers are the input to that argument, not the argument.
which they are before you argue for one; a single explanation asserted looks like the only one you thought of.
where they predict different things. This is often cheap once the main result exists — you already have the pipeline.
could refute is a story.
In the Discussion, and again in one sentence in the Results next to the effect it explains. A reader who sees the number and no account of it forms their own, and it is rarely yours.
Say that, and say what you would run. "We do not know why X and Y agree here; distinguishing the two accounts would need Z" is a real contribution and reads as honesty. Silence next to a surprising number reads as not having noticed it.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 39,894 | 30,435 | -24% | 1 | 1 | 0% | 2,270 | 2,634 | +16% | 0 | 0 | — |
case-02 | fail→fail | 17,388 | 16,275 | -6% | 1 | 1 | 0% | 2,480 | 2,556 | +3% | 0 | 0 | — |
case-03 | fail→fail | 16,253 | 16,385 | +1% | 1 | 1 | 0% | 2,363 | 2,447 | +4% | 0 | 0 | — |
case-04 | fail→pass | 26,702 | 14,713 | -45% | 1 | 1 | 0% | 1,685 | 2,306 | +37% | 0 | 0 | — |
case-05 | fail→pass | 12,356 | 10,162 | -18% | 1 | 1 | 0% | 1,798 | 1,689 | -6% | 0 | 0 | — |
case-06 | fail→pass | 23,121 | 33,054 | +43% | 1 | 1 | 0% | 1,837 | 2,223 | +21% | 0 | 0 | — |
case-07 | fail→fail | 7,008 | 29,824 | +326% | 1 | 1 | 0% | 930 | 2,033 | +119% | 0 | 0 | — |
case-08 | fail→pass | 12,801 | 13,703 | +7% | 1 | 1 | 0% | 2,156 | 2,356 | +9% | 0 | 0 | — |
case-09 | fail→pass | 31,630 | 35,720 | +13% | 1 | 1 | 0% | 1,771 | 2,767 | +56% | 0 | 0 | — |
case-10 | fail→pass | 16,440 | 17,547 | +7% | 1 | 1 | 0% | 2,203 | 2,671 | +21% | 0 | 0 | — |
case-11 | fail→pass | 19,307 | 17,974 | -7% | 1 | 1 | 0% | 2,553 | 2,636 | +3% | 0 | 0 | — |
case-12 | fail→pass | 10,174 | 50,907 | +400% | 1 | 1 | 0% | 1,512 | 2,033 | +34% | 0 | 0 | — |
case-13 | fail→fail | 16,176 | 12,770 | -21% | 1 | 1 | 0% | 1,779 | 1,789 | +1% | 0 | 0 | — |
case-14 | fail→pass | 14,354 | 16,728 | +17% | 1 | 1 | 0% | 1,811 | 2,403 | +33% | 0 | 0 | — |
case-15 | fail→pass | 18,159 | 14,255 | -21% | 1 | 1 | 0% | 2,354 | 2,266 | -4% | 0 | 0 | — |
case-16 | fail→pass | 12,042 | 9,806 | -19% | 1 | 1 | 0% | 1,449 | 1,808 | +25% | 0 | 0 | — |
case-17 | pass→pass | 28,700 | 9,778 | -66% | 1 | 1 | 0% | 1,646 | 1,529 | -7% | 0 | 0 | — |
case-18 | fail→pass | 18,988 | 19,925 | +5% | 1 | 1 | 0% | 2,697 | 2,974 | +10% | 0 | 0 | — |
case-19 | fail→pass | 15,294 | 13,284 | -13% | 1 | 1 | 0% | 2,096 | 2,247 | +7% | 0 | 0 | — |
case-20 | pass→pass | 13,544 | 12,284 | -9% | 1 | 1 | 0% | 1,377 | 1,382 | +0% | 0 | 0 | — |
case-21 | pass→fail | 40,297 | 30,888 | -23% | 1 | 1 | 0% | 2,810 | 3,080 | +10% | 0 | 0 | — |
case-22 | pass→pass | 10,745 | 12,305 | +15% | 1 | 1 | 0% | 1,547 | 2,149 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.