Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and analysis when a chemistry result is about to be reported in the units your code happens to produce, or without the program the field already uses. Covers anchoring to the incumbent, converting to the canonical unit, and turning an error distribution into a threshold success rate.
.claude/skills/tangxiangru-chemistry-canonical-units-thresholds-incumbent/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -8% | 0% |
A computational-chemistry number is only interpretable beside the program the field already runs, so budget baseline arms at design time: the established production code, the conventional architecture your method replaces, the explicit-parameter treatment your implicit one supersedes, and one deliberately cheap control (a geometric descriptor, an untrained network, a constant predictor) proving the added physics does work. Where the reference implementation is public and installable, install and run it on your systems instead of reimplementing it; a from-scratch reimplementation at a fraction of the original training budget cannot land on the published value, and only the same code path makes the comparison credible.
Report in the field's canonical unit with its canonical tolerance -- kcal/mol, meV/atom, eV/angstrom, angstrom, Debye, elementary charge -- and then convert error into a threshold success rate: fraction within chemical accuracy, fraction of poses under the accepted RMSD cut, fraction of systems inside the target band. A mean error alone hides the tail the field actually decides on.
Decompose every headline number along an axis the chemistry distinguishes and publish the decomposition, not just the pooled value: per dataset across the whole standard suite, per system or complex class, per charge and protonation state, per element, per interaction-distance regime, per functional group. Say how the rare, hard or imbalanced subset behaves. Repeat over seeds and splits with mean and standard deviation, and state the effective sample size when conformers, substitutions or targets cluster within families.
Chemistry's absent rate is only 12%; the losses come from unanchored and unstratified numbers. It fixes the measured 'method reported alone' failure (4 of 4 tasks demand an accepted reference on the same inputs), the pooled-aggregate failure (4 of 4 tasks demand a chemical decomposition and a threshold-conditioned rate rather than a mean), and the reimplementation trap -- the single highest-scoring criterion in the discipline (56) came from installing and running the published software, and mirroring the authors' released data outscored shipped-data-only runs 25.8 vs 15.4/18.3.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 19,526 | 16,712 | -14% | 1 | 1 | 0% | 3,217 | 2,962 | -8% | 0 | 0 | — |
case-02 | fail→pass | 23,619 | 22,357 | -5% | 1 | 1 | 0% | 3,407 | 3,828 | +12% | 0 | 0 | — |
case-03 | fail→pass | 34,458 | 23,461 | -32% | 1 | 1 | 0% | 5,040 | 4,305 | -15% | 0 | 0 | — |
case-04 | pass→pass | 21,965 | 16,181 | -26% | 1 | 1 | 0% | 3,466 | 3,011 | -13% | 0 | 0 | — |
case-05 | fail→pass | 59,978 | 43,718 | -27% | 1 | 1 | 0% | 3,194 | 3,432 | +7% | 0 | 0 | — |
case-06 | fail→pass | 15,111 | 11,746 | -22% | 1 | 1 | 0% | 2,433 | 2,247 | -8% | 0 | 0 | — |
case-07 | pass→pass | 31,803 | 20,039 | -37% | 1 | 1 | 0% | 3,331 | 3,519 | +6% | 0 | 0 | — |
case-08 | pass→pass | 31,761 | 13,744 | -57% | 1 | 1 | 0% | 2,557 | 2,660 | +4% | 0 | 0 | — |
case-09 | fail→pass | 18,914 | 15,369 | -19% | 1 | 1 | 0% | 2,526 | 2,787 | +10% | 0 | 0 | — |
case-10 | fail→pass | 25,840 | 14,940 | -42% | 1 | 1 | 0% | 2,690 | 2,691 | +0% | 0 | 0 | — |
case-11 | pass→pass | 18,572 | 29,973 | +61% | 1 | 1 | 0% | 2,960 | 3,219 | +9% | 0 | 0 | — |
case-12 | fail→pass | 17,242 | 31,753 | +84% | 1 | 1 | 0% | 2,345 | 2,441 | +4% | 0 | 0 | — |
case-13 | pass→pass | 26,576 | 15,535 | -42% | 1 | 1 | 0% | 1,886 | 2,520 | +34% | 0 | 0 | — |
case-14 | pass→pass | 46,247 | 24,989 | -46% | 1 | 1 | 0% | 2,700 | 3,263 | +21% | 0 | 0 | — |
case-15 | pass→pass | 31,781 | 24,682 | -22% | 1 | 1 | 0% | 3,090 | 3,294 | +7% | 0 | 0 | — |
case-16 | pass→pass | 18,946 | 19,557 | +3% | 1 | 1 | 0% | 2,748 | 3,444 | +25% | 0 | 0 | — |
case-17 | fail→pass | 16,965 | 43,769 | +158% | 1 | 1 | 0% | 2,715 | 2,466 | -9% | 0 | 0 | — |
case-18 | pass→pass | 20,359 | 21,532 | +6% | 1 | 1 | 0% | 2,565 | 2,243 | -13% | 0 | 0 | — |
case-19 | pass→pass | 24,085 | 40,319 | +67% | 1 | 1 | 0% | 2,868 | 2,784 | -3% | 0 | 0 | — |
case-20 | pass→pass | 16,645 | 22,581 | +36% | 1 | 1 | 0% | 2,572 | 2,613 | +2% | 0 | 0 | — |
case-21 | fail→fail | 22,087 | 64,889 | +194% | 1 | 1 | 0% | 3,258 | 4,445 | +36% | 0 | 0 | — |
case-22 | pass→pass | 44,762 | 46,484 | +4% | 1 | 1 | 0% | 3,349 | 4,216 | +26% | 0 | 0 | — |
case-23 | pass→fail | 19,757 | 26,984 | +37% | 1 | 1 | 0% | 3,002 | 4,921 | +64% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +35 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.