Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis and figure planning when a geospatial or gridded result is about to be reported only as regional aggregates. Covers reporting the stratified lattice, showing the field the strata came from, and which map a study of this kind is expected to publish.
.claude/skills/tangxiangru-earth-report-the-lattice-and-show-the-field/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 207% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 130% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 80% | 0% |
Before designing anything, read the archive's own layout: directory names, filename tokens, categorical columns, coordinate dimensions. Those are the axes the data distinguish, and in this field they are mandatory reporting axes rather than optional cuts. Deliver every headline quantity as a per-stratum table or small-multiple panel over the full cross-product the data support (method x sub-region, variable x level x lead, scenario x class, epoch x unit), each stratum carrying its own uncertainty, plus a total or global row. A pooled aggregate alone reads as the analysis never having been done.
Pair every skill, risk or change number with the spatial field that produced it: a map at the data's native resolution per scenario, period or lead time, plus zoomed regional panels over each place your result names, because coastal and mesoscale structure is invisible in a global panel and readers here look for it. Where the metric varies with lead time, horizon or epoch, plot it along that axis and mark the threshold crossing (skill horizon, sign change, acceleration), attaching an uncertainty to the change itself and not only to the levels.
Name and rank the specific regions, countries or basins that dominate the total, that carry the largest relative change, and where estimates disagree; a qualitative pattern description does not substitute. Print the governing numbers inside the panel. Produce the study's own stratified figure first; an original methodological question is an addition to it, never its replacement.
Twelve of Earth's seventeen criteria are image-typed and demand a decomposition rather than an aggregate. Earth_000 quantified every cross-method bias to within 0.01 of the paper and still scored 8 because its figure showed differences instead of per-region absolute values; Earth_003 scored 2 by showing 2 of 8 variables and 0 by mapping one lead time instead of a lead-time sequence. The one Earth image criterion that reached 50 did so because the numbers were printed inside the figure. The closing clause targets mechanism 2, novelty substitution: the two lowest-scoring Earth reports were the two most original.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 43,070 | 33,431 | -22% | 1 | 1 | 0% | 8,282 | 8,738 | +6% | 0 | 0 | — |
case-02 | fail→fail | 62,131 | 40,894 | -34% | 1 | 1 | 0% | 11,089 | 8,745 | -21% | 0 | 0 | — |
case-03 | fail→fail | 62,230 | 44,033 | -29% | 1 | 1 | 0% | 11,518 | 8,739 | -24% | 0 | 0 | — |
case-04 | pass→fail | 34,513 | 44,204 | +28% | 1 | 1 | 0% | 6,461 | 8,671 | +34% | 0 | 0 | — |
case-05 | pass→fail | 26,842 | 44,040 | +64% | 1 | 1 | 0% | 4,032 | 8,680 | +115% | 0 | 0 | — |
case-06 | pass→fail | 10,403 | 41,261 | +297% | 1 | 1 | 0% | 1,783 | 8,688 | +387% | 0 | 0 | — |
case-07 | fail→pass | 16,721 | 21,905 | +31% | 1 | 1 | 0% | 2,907 | 3,934 | +35% | 0 | 0 | — |
case-08 | fail→pass | 16,879 | 36,383 | +116% | 1 | 1 | 0% | 2,267 | 6,970 | +207% | 0 | 0 | — |
case-09 | pass→pass | 15,458 | 23,535 | +52% | 1 | 1 | 0% | 2,026 | 4,440 | +119% | 0 | 0 | — |
case-10 | fail→pass | 21,196 | 34,162 | +61% | 1 | 1 | 0% | 2,862 | 6,586 | +130% | 0 | 0 | — |
case-11 | fail→pass | 20,217 | 16,032 | -21% | 1 | 1 | 0% | 2,634 | 2,880 | +9% | 0 | 0 | — |
case-12 | pass→pass | 29,714 | 26,601 | -10% | 1 | 1 | 0% | 1,900 | 5,350 | +182% | 0 | 0 | — |
case-13 | fail→pass | 18,582 | 27,863 | +50% | 1 | 1 | 0% | 2,841 | 5,115 | +80% | 0 | 0 | — |
case-14 | fail→fail | 14,806 | 17,314 | +17% | 1 | 1 | 0% | 2,178 | 3,149 | +45% | 0 | 0 | — |
case-15 | fail→pass | 13,260 | 23,942 | +81% | 1 | 1 | 0% | 1,709 | 4,671 | +173% | 0 | 0 | — |
case-16 | pass→pass | 16,973 | 40,120 | +136% | 1 | 1 | 0% | 2,271 | 8,325 | +267% | 0 | 0 | — |
case-17 | pass→pass | 15,669 | 16,761 | +7% | 1 | 1 | 0% | 2,328 | 3,035 | +30% | 0 | 0 | — |
case-18 | fail→pass | 13,211 | 24,797 | +88% | 1 | 1 | 0% | 2,069 | 5,035 | +143% | 0 | 0 | — |
case-19 | pass→pass | 16,964 | 26,024 | +53% | 1 | 1 | 0% | 2,225 | 4,752 | +114% | 0 | 0 | — |
case-20 | fail→fail | 18,583 | 23,833 | +28% | 1 | 1 | 0% | 3,130 | 4,718 | +51% | 0 | 0 | — |
case-21 | fail→pass | 18,088 | 25,137 | +39% | 1 | 1 | 0% | 2,296 | 4,295 | +87% | 0 | 0 | — |
case-22 | fail→pass | 11,124 | 20,622 | +85% | 1 | 1 | 0% | 1,794 | 3,750 | +109% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.