Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at Stage 06 and Stage 07 when writing up a reproduction, and any time you have given a reproduced quantity, equation, figure or sequence a name of your own. Covers why a correct reproduction under private names reads as a missing one, which names have to be carried, and where they have to appear.
.claude/skills/tangxiangru-use-the-sources-own-names/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 32% | 0% |
You rebuilt the source's closed form and matched it on ninety-six of ninety-nine cases. In your report it is called the "GEO closed form", because that is what your module is named. The source calls it Equation 3. A reader looking for your verification of Equation 3 does not find one.
This is the cheapest loss in a reproduction and the hardest to see from inside, because everything is correct. The work was done, the numbers agree, the figure is there. Only the labels are yours, and labels are the entire interface between what you did and what anyone asked for.
For every object you reproduce, the source's name for it appears in your report, in the sentence that reports your result:
within 0.010 across 96 of 99 interfaces" — not "the GEO closed form agrees". Give your own name once, in parentheses, if you need it for the code.
paper's series is the Mackay sequence and its own new sequence is 1, 13, 45, 117, 239, 431, those digits belong in your text. A reader checking whether you reproduced the sequence looks for the sequence.
is one clause and it converts an unlabelled plot into a verification.
The comparators the task names are found by name or not at all.
redefined is a third quantity, and its agreement with the published one is a coincidence you have not checked.
There is usually a good reason the private name exists: your framing is more general, or your version fixed something, or the source's notation is bad. Keep it — after the source's name, in the same sentence. "Eq. (3) (which we implement in the more general form G, below)". The order matters because the first name is the one a reader matches against what they were looking for.
The same holds one level up. If your study reorganised the source's three results into your own five questions, the report still needs three headings a reader scanning for the source's results will land on. Your five questions can be the subsections under them.
Early. A reader forming a verdict on a figure often has the figure and the opening of the report, and not the section on page four where the symbol is defined. Put the source's name for the quantity, the equation number and the target value in the abstract and the first results paragraph — the same place the headline numbers go — and again in the figure caption and the axis label. A definition that arrives after the verdict has arrived too late.
Take the source's own list of results — its numbered equations, its figure captions, its abstract's claims. For each one this run reproduced, grep your report for the source's name for it. Every miss is a reproduction you performed and did not get credit for, and fixing it is a word.
See also reproduce-then-extend for the comparison table those names index, the-supplied-item-is-the-graded-unit for the identifier a shipped object keeps, and citation-discipline for pinning the reference the numbering belongs to.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,330 | 10,356 | -44% | 1 | 1 | 0% | 1,107 | 2,341 | +111% | 0 | 0 | — |
case-02 | fail→pass | 37,075 | 17,494 | -53% | 1 | 1 | 0% | 1,462 | 2,694 | +84% | 0 | 0 | — |
case-03 | fail→fail | 14,796 | 6,924 | -53% | 1 | 1 | 0% | 2,297 | 1,879 | -18% | 0 | 0 | — |
case-04 | fail→pass | 14,065 | 9,077 | -35% | 1 | 1 | 0% | 1,772 | 2,038 | +15% | 0 | 0 | — |
case-05 | pass→pass | 17,887 | 12,888 | -28% | 1 | 1 | 0% | 2,315 | 2,681 | +16% | 0 | 0 | — |
case-06 | fail→fail | 15,089 | 9,058 | -40% | 1 | 1 | 0% | 2,431 | 2,263 | -7% | 0 | 0 | — |
case-07 | pass→pass | 9,596 | 7,648 | -20% | 1 | 1 | 0% | 1,482 | 1,919 | +29% | 0 | 0 | — |
case-08 | fail→pass | 23,530 | 10,309 | -56% | 1 | 1 | 0% | 1,891 | 2,283 | +21% | 0 | 0 | — |
case-09 | fail→pass | 25,527 | 15,480 | -39% | 1 | 1 | 0% | 2,269 | 2,997 | +32% | 0 | 0 | — |
case-10 | fail→fail | 10,056 | 7,988 | -21% | 1 | 1 | 0% | 1,389 | 1,744 | +26% | 0 | 0 | — |
case-11 | pass→pass | 10,517 | 6,651 | -37% | 1 | 1 | 0% | 1,548 | 1,840 | +19% | 0 | 0 | — |
case-12 | fail→pass | 20,025 | 10,484 | -48% | 1 | 1 | 0% | 1,793 | 2,315 | +29% | 0 | 0 | — |
case-13 | pass→pass | 9,284 | 19,946 | +115% | 1 | 1 | 0% | 1,562 | 1,879 | +20% | 0 | 0 | — |
case-14 | pass→pass | 5,623 | 3,901 | -31% | 1 | 1 | 0% | 854 | 1,423 | +67% | 0 | 0 | — |
case-15 | fail→pass | 6,273 | 19,872 | +217% | 1 | 1 | 0% | 948 | 1,712 | +81% | 0 | 0 | — |
case-16 | pass→pass | 14,502 | 15,970 | +10% | 1 | 1 | 0% | 1,843 | 1,891 | +3% | 0 | 0 | — |
case-17 | pass→fail | 12,862 | 8,618 | -33% | 1 | 1 | 0% | 2,057 | 2,026 | -2% | 0 | 0 | — |
case-18 | fail→fail | 7,344 | 5,193 | -29% | 1 | 1 | 0% | 1,213 | 1,644 | +36% | 0 | 0 | — |
case-19 | pass→fail | 17,492 | 15,209 | -13% | 1 | 1 | 0% | 1,302 | 1,809 | +39% | 0 | 0 | — |
case-20 | pass→pass | 23,187 | 34,431 | +48% | 1 | 1 | 0% | 3,047 | 3,104 | +2% | 0 | 0 | — |
case-21 | pass→pass | 18,318 | 13,834 | -24% | 1 | 1 | 0% | 2,800 | 2,763 | -1% | 0 | 0 | — |
case-22 | pass→pass | 18,084 | 14,307 | -21% | 1 | 1 | 0% | 2,001 | 2,929 | +46% | 0 | 0 | — |
case-23 | pass→pass | 39,819 | 17,652 | -56% | 1 | 1 | 0% | 2,536 | 3,265 | +29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.