Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use whenever a number, a comparison or a claim is about to enter a stage summary or the report — at analysis and writing, and any time you are tempted to state a value you have not computed in this run. Covers where a number must come from, what to do when the experiment did not run, and why an honest gap outscores a plausible sentence.
.claude/skills/tangxiangru-evidence-not-assertion/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
A reader who cannot trace a number to an artifact has to decide whether to trust it, and a strict one decides no. That is not a style preference: a value you recall from the literature, infer from a trend, or round from memory is indistinguishable in the prose from one you measured, and the moment a single number turns out to be invented the whole report is worth nothing.
So, for every quantity that reaches the report:
code/ and written to a file under outputs/ orresults/ during this run. Name that file next to the number.
"The published value is X (Smith 2023); we measure Y" is a result. "The value is X" where X came from a paper is a fabrication with a citation missing.
Write what you did run and what remains unmeasured.
State it plainly, in one sentence, in the section where the result would have gone: what was not run, why, and what would settle it. Do not bury it in a limitations paragraph at the end, and do not substitute a proxy analysis without saying that is what you are doing.
An honest "we did not measure this" costs you that one result. A plausible sentence with no measurement behind it, once found, costs you the reader's belief in the results you did measure — including the good ones.
Take the three numbers your abstract leads with. For each, open the file it came from. If you cannot, it does not go in the abstract.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,714 | 7,348 | -37% | 1 | 1 | 0% | 1,390 | 1,571 | +13% | 0 | 0 | — |
case-02 | fail→pass | 8,213 | 7,649 | -7% | 1 | 1 | 0% | 1,240 | 1,376 | +11% | 0 | 0 | — |
case-03 | fail→pass | 8,141 | 5,579 | -31% | 1 | 1 | 0% | 1,152 | 1,238 | +7% | 0 | 0 | — |
case-04 | fail→pass | 8,313 | 6,657 | -20% | 1 | 1 | 0% | 1,192 | 1,152 | -3% | 0 | 0 | — |
case-05 | fail→pass | 8,963 | 7,407 | -17% | 1 | 1 | 0% | 1,277 | 1,368 | +7% | 0 | 0 | — |
case-06 | fail→pass | 37,441 | 13,073 | -65% | 1 | 1 | 0% | 1,808 | 1,643 | -9% | 0 | 0 | — |
case-07 | fail→pass | 20,065 | 5,474 | -73% | 1 | 1 | 0% | 1,405 | 1,211 | -14% | 0 | 0 | — |
case-08 | fail→pass | 15,179 | 19,683 | +30% | 1 | 1 | 0% | 2,133 | 2,036 | -5% | 0 | 0 | — |
case-09 | fail→pass | 8,847 | 5,829 | -34% | 1 | 1 | 0% | 1,264 | 1,276 | +1% | 0 | 0 | — |
case-10 | fail→pass | 14,457 | 6,493 | -55% | 1 | 1 | 0% | 1,314 | 1,345 | +2% | 0 | 0 | — |
case-11 | fail→pass | 12,898 | 6,580 | -49% | 1 | 1 | 0% | 1,670 | 1,157 | -31% | 0 | 0 | — |
case-12 | fail→pass | 15,130 | 5,847 | -61% | 1 | 1 | 0% | 1,557 | 1,262 | -19% | 0 | 0 | — |
case-13 | fail→pass | 9,220 | 15,420 | +67% | 1 | 1 | 0% | 1,359 | 1,411 | +4% | 0 | 0 | — |
case-14 | fail→pass | 9,984 | 7,091 | -29% | 1 | 1 | 0% | 1,449 | 1,370 | -5% | 0 | 0 | — |
case-15 | fail→pass | 10,906 | 20,932 | +92% | 1 | 1 | 0% | 1,376 | 1,240 | -10% | 0 | 0 | — |
case-16 | fail→pass | 14,335 | 10,133 | -29% | 1 | 1 | 0% | 2,124 | 1,701 | -20% | 0 | 0 | — |
case-17 | fail→pass | 9,796 | 5,716 | -42% | 1 | 1 | 0% | 1,325 | 1,123 | -15% | 0 | 0 | — |
case-18 | fail→pass | 8,242 | 7,456 | -10% | 1 | 1 | 0% | 1,238 | 1,405 | +13% | 0 | 0 | — |
case-19 | fail→pass | 10,188 | 7,656 | -25% | 1 | 1 | 0% | 1,432 | 1,410 | -2% | 0 | 0 | — |
case-20 | pass→pass | 9,771 | 12,114 | +24% | 1 | 1 | 0% | 1,385 | 1,459 | +5% | 0 | 0 | — |
case-21 | pass→pass | 4,523 | 5,918 | +31% | 1 | 1 | 0% | 653 | 1,237 | +89% | 0 | 0 | — |
case-22 | pass→pass | 12,151 | 30,515 | +151% | 1 | 1 | 0% | 2,006 | 3,018 | +50% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +86 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.