Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis when a materials result exists as a curve, a distribution or a trajectory and is about to be reported as one. Covers extracting the landmark scalar a reader compares — peak position, transition temperature, barrier height — in the property's physical unit, against a reference value.
.claude/skills/tangxiangru-material-landmark-scalars-in-physical-units/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 98% | 0% |
Materials results are judged as physical quantities, not as fit quality. For every curve, spectrum, surface or trace you compute, extract the landmarks the field quotes and print them as numbers: peak position and height, onset or threshold, fitted slope, equilibrium lattice parameter, transition temperature, iterations-to-target, yield above a property cut. Report accuracy in the property's own unit first -- meV/atom, eV, K, angstrom, GPa -- and only then any dimensionless goodness-of-fit. An R-squared, a rank correlation or a distributional divergence reported in place of a physical error reads as though the physical error was never measured. State the literature accuracy band for that property in the same unit and say whether you are inside it.
Every landmark needs an independent anchor computed on the same inputs: an experimental value, a higher-fidelity calculation, a published range, or a deliberately naive control (mean predictor, random selection, chance enrichment). Annotate the value, its uncertainty, N and the anchor directly inside the figure panel, and restate it in the abstract.
Never report only the pool. Give the number per material, per composition, per target value, per iteration budget, plus the aggregate over the full prescribed set, and check that the ordering across operating points matches the physically expected one. Separate interpolation from extrapolation explicitly: state where the training data's composition and property range ends, and report the error outside that range as its own number.
Addresses Material's two largest scoring mechanisms that are not absence: metric renaming (reporting R-squared/MMD where the field's quantity is an error in eV/atom or kelvin, which the judge reads as the demanded quantity being missing), and aggregate-instead-of-per-point reporting (three deep worked examples with no set-level statistic, a converged value with no iteration-indexed trace). It also forces the landmark out of prose into an annotated panel, which is where 10/13 image-typed Material criteria are read, and forces the separately-scored extrapolation/transfer number that 4 of 4 tasks demand.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 45,458 | 39,848 | -12% | 1 | 1 | 0% | 8,585 | 8,748 | +2% | 0 | 0 | — |
case-02 | fail→pass | 55,726 | 46,528 | -17% | 1 | 1 | 0% | 9,096 | 9,291 | +2% | 0 | 0 | — |
case-03 | fail→pass | 46,417 | 38,188 | -18% | 1 | 1 | 0% | 8,368 | 8,261 | -1% | 0 | 0 | — |
case-04 | fail→pass | 16,092 | 17,271 | +7% | 1 | 1 | 0% | 2,460 | 3,606 | +47% | 0 | 0 | — |
case-05 | fail→fail | 32,233 | 101,485 | +215% | 1 | 1 | 0% | 4,600 | 6,659 | +45% | 0 | 0 | — |
case-06 | fail→pass | 23,010 | 35,756 | +55% | 1 | 1 | 0% | 3,695 | 7,326 | +98% | 0 | 0 | — |
case-07 | fail→fail | 46,327 | 21,672 | -53% | 1 | 1 | 0% | 1,234 | 4,013 | +225% | 0 | 0 | — |
case-08 | fail→pass | 16,994 | 34,768 | +105% | 1 | 1 | 0% | 2,856 | 7,385 | +159% | 0 | 0 | — |
case-09 | fail→pass | 31,753 | 31,635 | -0% | 1 | 1 | 0% | 1,255 | 6,012 | +379% | 0 | 0 | — |
case-10 | fail→pass | 39,761 | 56,864 | +43% | 1 | 1 | 0% | 3,732 | 8,690 | +133% | 0 | 0 | — |
case-11 | fail→fail | 40,662 | 25,901 | -36% | 1 | 1 | 0% | 2,138 | 5,122 | +140% | 0 | 0 | — |
case-12 | fail→pass | 13,892 | 23,748 | +71% | 1 | 1 | 0% | 2,123 | 4,739 | +123% | 0 | 0 | — |
case-13 | fail→pass | 23,111 | 73,342 | +217% | 1 | 1 | 0% | 3,834 | 8,412 | +119% | 0 | 0 | — |
case-14 | fail→pass | 27,054 | 45,115 | +67% | 1 | 1 | 0% | 2,817 | 7,339 | +161% | 0 | 0 | — |
case-15 | fail→pass | 21,659 | 43,619 | +101% | 1 | 1 | 0% | 4,473 | 8,691 | +94% | 0 | 0 | — |
case-16 | fail→pass | 19,278 | 43,331 | +125% | 1 | 1 | 0% | 2,884 | 8,440 | +193% | 0 | 0 | — |
case-17 | fail→fail | 19,789 | 28,970 | +46% | 1 | 1 | 0% | 3,134 | 5,780 | +84% | 0 | 0 | — |
case-18 | pass→pass | 10,893 | 24,344 | +123% | 1 | 1 | 0% | 1,895 | 4,693 | +148% | 0 | 0 | — |
case-19 | pass→fail | 22,370 | 76,124 | +240% | 1 | 1 | 0% | 4,115 | 8,704 | +112% | 0 | 0 | — |
case-20 | pass→fail | 21,685 | 49,944 | +130% | 1 | 1 | 0% | 3,958 | 7,257 | +83% | 0 | 0 | — |
case-21 | fail→pass | 68,831 | 66,332 | -4% | 1 | 1 | 0% | 4,599 | 7,495 | +63% | 0 | 0 | — |
case-22 | fail→fail | 21,062 | 68,722 | +226% | 1 | 1 | 0% | 3,096 | 7,181 | +132% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.