Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis and figure planning when the computation ranks entities — molecules, poses, fragments, atoms — or sweeps a property along a coordinate. Covers printing the named ranked list and the property-versus-coordinate curve, the two artifacts most often computed here and least often reported.
.claude/skills/tangxiangru-chemistry-ranked-entities-and-property-curves/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 106% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -40% | 0% |
Two chemistry deliverables are routinely computed and never printed. Plan both into the report skeleton.
First, the ranked list. Whenever the method scores individual entities -- per-residue scans, per-atom or per-fragment attributions, per-pose scores, per-molecule rankings -- the deliverable is an explicit table of the leading entities by chemical identifier with their scores and units, plus the scan's bookkeeping: how many entities were scanned and the observed minimum and maximum. An aggregate ranking metric (AUC, precision@k, a correlation) does not substitute for the named list. If a per-entity file exists in your outputs, sorting its head into the report is the result. Then map those entities back onto chemistry -- which contacts, which functional groups, which charges or multipoles -- and state whether that is what a chemist would expect.
Second, the curve. Where a property depends on a governing physical or protocol coordinate -- bond length, intermolecular separation, cutoff radius, temperature, training-set size, number of sampling steps -- sweep it and plot the continuous curve with the reference overlaid, then read the derived constants off it and report their errors: equilibrium geometry, well depth, barrier height, asymptotic decay exponent, sum rules and conservation checks. A bar chart of RMSE by model variant answers a different question and does not replace it.
Make both artifacts self-supporting: metric value, N and the comparison annotated inside the panel, and the headline restated in the abstract in the field's units.
These are the two documented computed-but-never-printed failures. One run's per-entity output file already held the exact ranked list and range the hidden criterion wanted; the report published a pooled AUC and precision@k instead and scored 18 and 0 across two runs -- one sort-and-head away from the answer. Separately, a right result delivered as a bar chart of RMSE by variant instead of the energy-versus-separation curve with the reference overlaid scored 45, 5, 5 across three runs. It also forces the interpretation-onto-named-chemical-entities demand that appears in 4 of 4 tasks.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,988 | 40,779 | +46% | 1 | 1 | 0% | 4,534 | 7,429 | +64% | 0 | 0 | — |
case-02 | fail→pass | 24,978 | 31,715 | +27% | 1 | 1 | 0% | 4,280 | 5,803 | +36% | 0 | 0 | — |
case-03 | fail→pass | 24,159 | 44,546 | +84% | 1 | 1 | 0% | 4,248 | 8,751 | +106% | 0 | 0 | — |
case-04 | fail→fail | 19,501 | 15,915 | -18% | 1 | 1 | 0% | 2,821 | 2,816 | -0% | 0 | 0 | — |
case-05 | fail→fail | 13,306 | 15,980 | +20% | 1 | 1 | 0% | 1,799 | 2,928 | +63% | 0 | 0 | — |
case-06 | fail→pass | 16,338 | 17,542 | +7% | 1 | 1 | 0% | 2,607 | 2,880 | +10% | 0 | 0 | — |
case-07 | fail→pass | 16,411 | 15,791 | -4% | 1 | 1 | 0% | 2,222 | 2,670 | +20% | 0 | 0 | — |
case-08 | pass→pass | 16,120 | 12,877 | -20% | 1 | 1 | 0% | 2,336 | 2,354 | +1% | 0 | 0 | — |
case-09 | fail→fail | 18,515 | 14,792 | -20% | 1 | 1 | 0% | 2,508 | 2,620 | +4% | 0 | 0 | — |
case-10 | fail→fail | 14,167 | 18,006 | +27% | 1 | 1 | 0% | 2,113 | 3,069 | +45% | 0 | 0 | — |
case-11 | fail→fail | 13,951 | 15,259 | +9% | 1 | 1 | 0% | 2,070 | 2,965 | +43% | 0 | 0 | — |
case-12 | fail→fail | 20,078 | 15,916 | -21% | 1 | 1 | 0% | 2,638 | 2,591 | -2% | 0 | 0 | — |
case-13 | fail→pass | 15,891 | 7,137 | -55% | 1 | 1 | 0% | 2,395 | 1,430 | -40% | 0 | 0 | — |
case-14 | fail→pass | 6,581 | 9,917 | +51% | 1 | 1 | 0% | 948 | 1,829 | +93% | 0 | 0 | — |
case-15 | fail→fail | 16,250 | 14,901 | -8% | 1 | 1 | 0% | 2,335 | 2,740 | +17% | 0 | 0 | — |
case-16 | fail→fail | 16,806 | 5,981 | -64% | 1 | 1 | 0% | 2,499 | 1,242 | -50% | 0 | 0 | — |
case-17 | fail→fail | 13,299 | 13,888 | +4% | 1 | 1 | 0% | 2,101 | 2,453 | +17% | 0 | 0 | — |
case-18 | fail→fail | 15,416 | 11,039 | -28% | 1 | 1 | 0% | 2,138 | 2,100 | -2% | 0 | 0 | — |
case-19 | fail→fail | 15,553 | 8,403 | -46% | 1 | 1 | 0% | 2,316 | 1,590 | -31% | 0 | 0 | — |
case-20 | fail→pass | 17,217 | 21,982 | +28% | 1 | 1 | 0% | 2,664 | 3,566 | +34% | 0 | 0 | — |
case-21 | pass→pass | 33,043 | 39,038 | +18% | 1 | 1 | 0% | 2,692 | 7,191 | +167% | 0 | 0 | — |
case-22 | pass→pass | 27,636 | 44,703 | +62% | 1 | 1 | 0% | 4,521 | 8,726 | +93% | 0 | 0 | — |
case-23 | pass→pass | 15,025 | 45,966 | +206% | 1 | 1 | 0% | 2,344 | 8,709 | +272% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +30 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.