Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at Stage 05 and Stage 06 the moment a reproduction lands materially off a number the source study published — a different order of magnitude, an inverted trend, a collapsed estimate. Covers why the gap is a defect in your pipeline until you have shown otherwise, how much of the remaining budget to spend closing it, and what to write when it will not close.
.claude/skills/tangxiangru-close-the-gap-to-the-published-number/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 5% | 0% |
You reran the study and got 0.140 V where the paper got 0.0117 V. You have two moves available. One is to write the discrepancy down carefully, attribute it to a difference in setup, and move on to the next section. The other is to treat the gap as evidence that something in your pipeline is wrong, and spend budget until it closes or until you can name the specific thing that makes it irreducible.
The first move feels like honesty and reads as one. It is also, almost always, wrong on the facts: a twelvefold gap against a published measurement is far more often a bug than a boundary. And it is scored as what it is — an analysis whose methodology is defensible and whose numbers are not.
A quantitative disagreement with the source is a defect with an owner until an experiment says otherwise. Not a limitation, not a caveat, not a scope note. The owner is this run.
Order of magnitude matters. Treat these differently:
| gap | reading | |---|---| | within the source's own stated uncertainty | agreement; say so and move on | | a factor of 1.5-3, or a shifted intercept | a parameter, a normalisation, a split, a unit | | a factor of ten, an inverted sign, an inverted trend | a bug; nothing else does this | | your estimator collapses to a constant or to zero | a bug; the model learnt nothing |
The last two rows are not results. A latent charge that collapses to 0.0002 e, a curve that sits exactly on the best-constant null, a metric that declines where the paper's improves — each of these is a pipeline that did not run, wearing the clothes of a finding.
Before you write the discrepancy up, run the cheap discriminators. Most gaps die to one of them:
versus MSE, percentage versus fraction. Recompute the published number in your units by hand, once, on paper.
set, or a curated subset? Are you scoring the same rows?
that trained on thousands is not a reproduction of the method, it is a measurement of your sample size. Say what you trained on and how far under the source you are, and if the shortfall is affordable, fix it rather than reporting it.
schedule that is silently wrong will burn a hundred training runs and will not announce itself. Check the ones you invented, first.
trivially checkable answer and reproduce that. If it fails, the gap is upstream of the science.
Do this while the compute budget is still open. A gap discovered at Stage 06 with nothing left to run is a gap you will describe rather than close; the time to look is the hour after the first number lands, not the hour before writing.
Then it becomes a finding, and it is owed the same rigour as any other:
"Our number differs and we discuss possible reasons" is not that. A list of plausible causes with no eliminations behind it reads as a run that did not look.
For every quantity the source publishes and this task asks for, one of three is true and written down: it agrees, it disagrees and you closed it, or it disagrees and you eliminated the cheap causes and named the expensive one. Nothing in the fourth state — noticed, described, unexamined.
See also reproduce-then-extend for how the comparison table is built and where the published column comes from, and run-the-requested-analysis for what to do when the gap turns out to be in the supplied inputs rather than in your code.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 26,833 | 31,039 | +16% | 1 | 1 | 0% | 4,323 | 4,535 | +5% | 0 | 0 | — |
case-02 | fail→pass | 29,514 | 26,405 | -11% | 1 | 1 | 0% | 3,390 | 3,250 | -4% | 0 | 0 | — |
case-03 | fail→pass | 32,187 | 14,178 | -56% | 1 | 1 | 0% | 2,821 | 2,820 | -0% | 0 | 0 | — |
case-04 | fail→fail | 26,286 | 14,187 | -46% | 1 | 1 | 0% | 2,846 | 2,835 | -0% | 0 | 0 | — |
case-05 | pass→pass | 16,353 | 11,250 | -31% | 1 | 1 | 0% | 2,332 | 2,193 | -6% | 0 | 0 | — |
case-06 | fail→pass | 15,719 | 8,582 | -45% | 1 | 1 | 0% | 2,548 | 2,099 | -18% | 0 | 0 | — |
case-07 | pass→pass | 18,125 | 13,511 | -25% | 1 | 1 | 0% | 2,402 | 3,151 | +31% | 0 | 0 | — |
case-08 | pass→pass | 19,158 | 11,433 | -40% | 1 | 1 | 0% | 2,074 | 2,606 | +26% | 0 | 0 | — |
case-09 | pass→pass | 23,594 | 22,781 | -3% | 1 | 1 | 0% | 2,246 | 2,032 | -10% | 0 | 0 | — |
case-10 | pass→pass | 17,659 | 18,464 | +5% | 1 | 1 | 0% | 2,788 | 3,873 | +39% | 0 | 0 | — |
case-11 | pass→pass | 15,649 | 5,140 | -67% | 1 | 1 | 0% | 2,688 | 1,869 | -30% | 0 | 0 | — |
case-12 | pass→pass | 58,358 | 16,030 | -73% | 1 | 1 | 0% | 3,510 | 3,080 | -12% | 0 | 0 | — |
case-13 | pass→pass | 14,900 | 17,603 | +18% | 1 | 1 | 0% | 1,900 | 1,999 | +5% | 0 | 0 | — |
case-14 | pass→pass | 17,375 | 10,331 | -41% | 1 | 1 | 0% | 2,521 | 2,417 | -4% | 0 | 0 | — |
case-15 | pass→pass | 25,098 | 34,948 | +39% | 1 | 1 | 0% | 1,680 | 2,011 | +20% | 0 | 0 | — |
case-16 | fail→pass | 15,064 | 12,021 | -20% | 1 | 1 | 0% | 2,202 | 2,458 | +12% | 0 | 0 | — |
case-17 | pass→pass | 12,126 | 9,858 | -19% | 1 | 1 | 0% | 1,635 | 1,826 | +12% | 0 | 0 | — |
case-18 | pass→pass | 18,779 | 11,029 | -41% | 1 | 1 | 0% | 2,691 | 2,274 | -15% | 0 | 0 | — |
case-19 | fail→pass | 13,877 | 7,614 | -45% | 1 | 1 | 0% | 1,860 | 1,959 | +5% | 0 | 0 | — |
case-20 | pass→pass | 14,645 | 7,848 | -46% | 1 | 1 | 0% | 2,064 | 2,055 | -0% | 0 | 0 | — |
case-21 | pass→pass | 30,248 | 45,184 | +49% | 1 | 1 | 0% | 2,477 | 2,699 | +9% | 0 | 0 | — |
case-22 | pass→pass | 29,995 | 17,560 | -41% | 1 | 1 | 0% | 3,380 | 3,412 | +1% | 0 | 0 | — |
case-23 | pass→pass | 19,086 | 19,729 | +3% | 1 | 1 | 0% | 2,754 | 2,686 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.