Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the task is to reproduce, replicate or re-implement a published study, at design time and when reporting results. Covers what a reproduction must report, how to compare against the source study's numbers, and why the reproduction comes before any improvement.
.claude/skills/tangxiangru-reproduce-then-extend/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 2% | 0% |
When the task is to reproduce published work, the deliverable is not "a good study of this topic". It is: the same quantities, measured the same way, reported next to the published ones, with the agreement or disagreement stated.
Build it early and keep it visible:
| quantity | published | ours | agree? |
One row per quantity the source study reports and this task asks for. Fill the published column at the literature stage, before you have your own numbers — that ordering is what keeps the comparison honest, because a target read after the fact tends to become the answer you find.
Where you cannot recompute a published value, say so in the row rather than leaving it blank. Where your number disagrees, say by how much and give your best account of why; a disagreement you explain is a finding, and one you omit is an error the reader will find for you.
If you see a better method, the order still matters: reproduce what the paper did, report it, and then show your improvement as a separate result against your own reproduction. An improvement reported without the reproduction leaves the reader unable to tell whether you beat the paper or measured something else.
Reproduce every arm the task names, not the one that worked. A reproduction that covers one of three models is a reproduction of one third of the study, and the missing two are what the reader notices.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,615 | 17,461 | -6% | 1 | 1 | 0% | 2,896 | 2,864 | -1% | 0 | 0 | — |
case-02 | fail→fail | 30,170 | 25,851 | -14% | 1 | 1 | 0% | 1,026 | 4,232 | +312% | 0 | 0 | — |
case-03 | fail→fail | 43,095 | 19,275 | -55% | 1 | 1 | 0% | 3,229 | 3,631 | +12% | 0 | 0 | — |
case-04 | pass→pass | 28,363 | 31,571 | +11% | 1 | 1 | 0% | 4,639 | 4,350 | -6% | 0 | 0 | — |
case-05 | pass→pass | 41,877 | 22,715 | -46% | 1 | 1 | 0% | 3,213 | 3,680 | +15% | 0 | 0 | — |
case-06 | pass→pass | 15,560 | 13,378 | -14% | 1 | 1 | 0% | 2,238 | 2,082 | -7% | 0 | 0 | — |
case-07 | pass→fail | 15,728 | 17,125 | +9% | 1 | 1 | 0% | 2,164 | 2,720 | +26% | 0 | 0 | — |
case-08 | fail→pass | 16,853 | 12,015 | -29% | 1 | 1 | 0% | 2,346 | 2,166 | -8% | 0 | 0 | — |
case-09 | fail→pass | 17,065 | 12,432 | -27% | 1 | 1 | 0% | 2,352 | 2,262 | -4% | 0 | 0 | — |
case-10 | fail→pass | 19,204 | 15,131 | -21% | 1 | 1 | 0% | 2,560 | 2,489 | -3% | 0 | 0 | — |
case-11 | pass→pass | 14,286 | 13,505 | -5% | 1 | 1 | 0% | 1,828 | 2,090 | +14% | 0 | 0 | — |
case-12 | pass→pass | 14,311 | 14,039 | -2% | 1 | 1 | 0% | 1,830 | 2,143 | +17% | 0 | 0 | — |
case-13 | fail→pass | 14,270 | 27,822 | +95% | 1 | 1 | 0% | 2,352 | 2,393 | +2% | 0 | 0 | — |
case-14 | fail→pass | 14,672 | 10,117 | -31% | 1 | 1 | 0% | 2,120 | 1,927 | -9% | 0 | 0 | — |
case-15 | fail→pass | 17,017 | 13,458 | -21% | 1 | 1 | 0% | 2,388 | 2,263 | -5% | 0 | 0 | — |
case-16 | pass→pass | 12,305 | 11,493 | -7% | 1 | 1 | 0% | 1,906 | 2,069 | +9% | 0 | 0 | — |
case-17 | pass→pass | 10,923 | 10,389 | -5% | 1 | 1 | 0% | 1,348 | 1,866 | +38% | 0 | 0 | — |
case-18 | fail→pass | 11,900 | 8,192 | -31% | 1 | 1 | 0% | 1,667 | 1,539 | -8% | 0 | 0 | — |
case-19 | pass→pass | 18,428 | 11,173 | -39% | 1 | 1 | 0% | 2,134 | 1,953 | -8% | 0 | 0 | — |
case-20 | pass→pass | 14,407 | 9,735 | -32% | 1 | 1 | 0% | 1,889 | 1,702 | -10% | 0 | 0 | — |
case-21 | fail→pass | 14,444 | 11,925 | -17% | 1 | 1 | 0% | 2,086 | 2,097 | +1% | 0 | 0 | — |
case-22 | pass→pass | 15,950 | 10,951 | -31% | 1 | 1 | 0% | 2,245 | 1,882 | -16% | 0 | 0 | — |
case-23 | fail→pass | 11,423 | 3,893 | -66% | 1 | 1 | 0% | 1,597 | 824 | -48% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.