Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design when planning figures for a reproduction, replication or validation task, and again before the report is written. Covers deriving each panel's series list and axis ranges from the source's rendered figure, giving every source result a panel before your own hypotheses claim the slots, and printing the source's named constants as labelled values.
.claude/skills/tangxiangru-draw-the-source-figure-panel-for-panel/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 7% | 0% |
A reproduction is read by putting the source's figure beside yours and asking whether they show the same thing. That is what a reader does with a reproduction, and it is the only form in which "we got the same answer" is checkable. Your study design is not the index anyone reads you under: a result you settled early, or discharged with a headline number, still needs its panel.
A run reproduced a packing theory to 110 of 111 published closed-form values — more of the paper reproduced than either comparator managed — and scored 18.7 where a plain agent that did less scored 53.4. Every figure slot in its plan was bound to one of its own preregistered hypotheses, and the paper's first result was discharged as the headline number "110 of 111". That result is a lattice-path construction carrying two magic-number series. Both were on disk. Neither was drawn. The word "magic" appears once in the entire report, in a bibliography entry, and the requirement scored 20 against 55 and 48 for the two runs that drew it.
The mechanism is not how many panels exist — that run had spare slots, and other runs in the same batch shipped fourteen panels and still lost the comparison. The mechanism is which result owns a panel. An agreement count and a CDF are a better summary and a worse exhibit than the panel they replace.
A second shape: a scan panel carried two series where the published panel carries three, because the run reproduced a caption it had extracted from an older preprint rather than the rendered figure. The third series was 250 rows on disk and was already plotted in the neighbouring panel.
notes/source_figures.json: one row per panel ofevery figure the task points at — source figure and panel letter, x and y quantities, the full list of series, axis ranges, special markers (an inferred point drawn open, a shaded band, a hardness line) and the values printed on it. Read these off the rendered figure. Extracted captions come from whichever version you fetched and routinely disagree with the published panel about how many series it carries.
plan schema demands a claim id, use exploratory:reproduce-source-figure-<n>. A reproduction panel is not exploratory in any sense that matters.
three-panel figure with the panels in that order, so the two can be laid side by side without the reader doing translation.
published 0.04 / ours 0.0400 — with any agreement count beside them, never instead of them. A count says how often you agreed; the labelled value is the agreement.
Deleting it deletes the comparison; annotating it adds a finding and keeps both.
still gets a panel: the source's value against your prediction, labelled as a proxy.
section, early. A figure whose supporting sentence arrives forty thousand characters into the report is read on the picture alone.
Walk the panel list. For each row, name your file and panel and diff the series lists against the source's. Any row whose answer is "a table", "a headline number" or "prose" is a result you did not draw.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 103,537 | 29,731 | -71% | 1 | 1 | 0% | 3,627 | 3,747 | +3% | 0 | 0 | — |
case-02 | fail→pass | 26,698 | 10,676 | -60% | 1 | 1 | 0% | 2,556 | 2,373 | -7% | 0 | 0 | — |
case-03 | fail→pass | 27,486 | 21,868 | -20% | 1 | 1 | 0% | 2,913 | 3,659 | +26% | 0 | 0 | — |
case-04 | pass→pass | 16,606 | 9,786 | -41% | 1 | 1 | 0% | 1,977 | 2,399 | +21% | 0 | 0 | — |
case-05 | fail→pass | 14,553 | 3,911 | -73% | 1 | 1 | 0% | 1,884 | 1,415 | -25% | 0 | 0 | — |
case-06 | pass→pass | 17,883 | 10,680 | -40% | 1 | 1 | 0% | 2,340 | 2,472 | +6% | 0 | 0 | — |
case-07 | pass→pass | 28,402 | 8,159 | -71% | 1 | 1 | 0% | 1,753 | 2,031 | +16% | 0 | 0 | — |
case-08 | pass→pass | 25,237 | 5,714 | -77% | 1 | 1 | 0% | 1,662 | 1,694 | +2% | 0 | 0 | — |
case-09 | pass→pass | 14,747 | 5,720 | -61% | 1 | 1 | 0% | 1,955 | 1,669 | -15% | 0 | 0 | — |
case-10 | fail→pass | 32,026 | 6,503 | -80% | 1 | 1 | 0% | 1,840 | 1,612 | -12% | 0 | 0 | — |
case-11 | fail→pass | 13,152 | 8,875 | -33% | 1 | 1 | 0% | 1,723 | 1,851 | +7% | 0 | 0 | — |
case-12 | fail→pass | 14,644 | 6,711 | -54% | 1 | 1 | 0% | 1,985 | 1,403 | -29% | 0 | 0 | — |
case-13 | fail→pass | 11,395 | 6,051 | -47% | 1 | 1 | 0% | 1,616 | 1,636 | +1% | 0 | 0 | — |
case-14 | fail→pass | 35,655 | 20,329 | -43% | 1 | 1 | 0% | 2,679 | 2,332 | -13% | 0 | 0 | — |
case-15 | fail→pass | 17,472 | 4,673 | -73% | 1 | 1 | 0% | 2,276 | 1,198 | -47% | 0 | 0 | — |
case-16 | fail→pass | 12,619 | 6,288 | -50% | 1 | 1 | 0% | 1,852 | 1,734 | -6% | 0 | 0 | — |
case-17 | pass→pass | 10,397 | 4,932 | -53% | 1 | 1 | 0% | 1,385 | 1,542 | +11% | 0 | 0 | — |
case-18 | fail→pass | 23,057 | 58,601 | +154% | 1 | 1 | 0% | 3,071 | 3,526 | +15% | 0 | 0 | — |
case-19 | fail→fail | 12,826 | 7,910 | -38% | 1 | 1 | 0% | 1,614 | 1,918 | +19% | 0 | 0 | — |
case-20 | pass→pass | 23,807 | 24,440 | +3% | 1 | 1 | 0% | 3,600 | 4,275 | +19% | 0 | 0 | — |
case-21 | pass→pass | 23,970 | 26,107 | +9% | 1 | 1 | 0% | 2,497 | 2,955 | +18% | 0 | 0 | — |
case-22 | pass→pass | 23,641 | 25,526 | +8% | 1 | 1 | 0% | 3,377 | 4,451 | +32% | 0 | 0 | — |
case-23 | pass→pass | 23,454 | 39,232 | +67% | 1 | 1 | 0% | 2,947 | 6,681 | +127% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +48 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.