Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design when choosing the figure list, and again before writing, when a result is about to be reported as a pooled number or a table. Covers why the figure is the unit a result is delivered in here, which panels a paper of this kind is expected to carry, and what a pooled number hides.
.claude/skills/tangxiangru-astronomy-figure-is-the-unit-of-result/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 2725% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 19% | 0% |
Draft your figure list before your analysis list. In observational astronomy the result is delivered as a figure a reader can lay beside the published one: the same quantities on both axes, the same log or linear scaling, the same decade span, the full family of competing models or method variants overlaid on one set of axes, the decision threshold drawn as a labelled line, and a caption stating the numbers a reader would otherwise have to read off the plot (N, median, confidence level, how many points fall each side of the line). A result that is scientifically equivalent but visually reorganised, or that exists only as a table or a sentence, will not be recognised as that result.
Nothing is reported only as an aggregate. Break every quantity down by the sub-population the data physically distinguish, per object, per source class, per component index, per bin of the independent variable, quantify the trend across that index, and say which component dominates the budget and by what factor. Where several independent indicators, models or approximations exist for one quantity, put them all on a single axis and quantify their disagreement; where two independent routes measure the same object, state the consistency with a significance.
Present the result as a curve over the scanned variable across its full plausible range, log axes when it spans decades, and state the interval over which the claim actually holds, rather than a value at one point.
Seven of Astronomy's twelve criteria are image-typed and carry 60% of the weight, judged by a vision model that is shown the original figure first. The measured mechanism 'right numbers, wrong artifact class' cost a whole criterion: the agent had the residual information in prose and in a differently-shaped bar chart and scored near zero because the demanded panel was never drawn. The decomposition clause fixes the recurring aggregate-only loss, and the single-axis multi-indicator clause fixes the case where the canonical multi-model comparison figure was never produced (score 0) because the agent's own question did not need it.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 95,115 | 108,579 | +14% | 1 | 1 | 0% | 3,908 | 8,665 | +122% | 0 | 0 | — |
case-02 | fail→pass | 11,114 | 45,588 | +310% | 1 | 1 | 0% | 315 | 8,900 | +2725% | 0 | 0 | — |
case-03 | pass→pass | 14,746 | 20,868 | +42% | 1 | 1 | 0% | 2,159 | 3,742 | +73% | 0 | 0 | — |
case-04 | fail→fail | 90,306 | 158,730 | +76% | 1 | 1 | 0% | 7,787 | 8,974 | +15% | 0 | 0 | — |
case-05 | fail→pass | 41,821 | 149,686 | +258% | 1 | 1 | 0% | 7,558 | 8,552 | +13% | 0 | 0 | — |
case-06 | fail→pass | 24,807 | 34,366 | +39% | 1 | 1 | 0% | 3,855 | 6,242 | +62% | 0 | 0 | — |
case-07 | fail→fail | 62,113 | 48,603 | -22% | 1 | 1 | 0% | 4,293 | 8,673 | +102% | 0 | 0 | — |
case-08 | fail→pass | 21,503 | 55,401 | +158% | 1 | 1 | 0% | 2,463 | 4,247 | +72% | 0 | 0 | — |
case-09 | fail→pass | 23,489 | 26,619 | +13% | 1 | 1 | 0% | 3,573 | 4,267 | +19% | 0 | 0 | — |
case-10 | fail→pass | 23,012 | 26,186 | +14% | 1 | 1 | 0% | 3,085 | 4,289 | +39% | 0 | 0 | — |
case-11 | fail→pass | 5,086 | 18,051 | +255% | 1 | 1 | 0% | 641 | 3,243 | +406% | 0 | 0 | — |
case-12 | fail→pass | 19,639 | 27,716 | +41% | 1 | 1 | 0% | 3,051 | 4,985 | +63% | 0 | 0 | — |
case-13 | fail→pass | 19,207 | 41,433 | +116% | 1 | 1 | 0% | 1,220 | 5,393 | +342% | 0 | 0 | — |
case-14 | fail→pass | 18,664 | 28,679 | +54% | 1 | 1 | 0% | 1,253 | 5,320 | +325% | 0 | 0 | — |
case-15 | fail→pass | 14,851 | 50,240 | +238% | 1 | 1 | 0% | 2,010 | 4,796 | +139% | 0 | 0 | — |
case-16 | fail→pass | 21,816 | 35,805 | +64% | 1 | 1 | 0% | 3,532 | 3,790 | +7% | 0 | 0 | — |
case-17 | fail→pass | 16,482 | 92,845 | +463% | 1 | 1 | 0% | 2,500 | 6,159 | +146% | 0 | 0 | — |
case-18 | fail→pass | 31,331 | 21,026 | -33% | 1 | 1 | 0% | 2,465 | 4,070 | +65% | 0 | 0 | — |
case-19 | fail→fail | 7,869 | 17,911 | +128% | 1 | 1 | 0% | 1,145 | 2,933 | +156% | 0 | 0 | — |
case-20 | fail→pass | 10,614 | 27,925 | +163% | 1 | 1 | 0% | 1,450 | 4,596 | +217% | 0 | 0 | — |
case-21 | pass→pass | 20,961 | 27,758 | +32% | 1 | 1 | 0% | 3,698 | 4,875 | +32% | 0 | 0 | — |
case-22 | fail→pass | 13,547 | 101,004 | +646% | 1 | 1 | 0% | 2,148 | 7,750 | +261% | 0 | 0 | — |
case-23 | fail→pass | 20,885 | 37,923 | +82% | 1 | 1 | 0% | 3,803 | 7,011 | +84% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +74 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.