Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at Stage 06 and Stage 07 when a deliverable the task named cannot be produced by this run at all — a wet-lab measurement, a synthesised material, a proprietary benchmark, hardware you do not have. Covers the difference between fabricating a number and citing one, where the cited value belongs, and why omitting the section is the worst of the three options.
.claude/skills/tangxiangru-a-value-you-did-not-measure-still-has-a-source/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 26% | 0% |
Some tasks name a deliverable this run cannot produce. The study asks for a synthesised polymer and a DSC trace; you have no bench. It asks for a run on hardware you do not have, or a benchmark behind a licence, or an experiment that takes six weeks.
There are three things you can do, and they are not close in value:
from a run that forgot, and it is scored as absence.
here is why. Honest, and better than silence.
it. The value exists — in the source study, in a reference database, in a handbook. Give it, say whose it is, and show what your work does and does not establish about it.
The third is what a real paper does. A methods section that needs a melting point it did not measure cites one. A benchmark table with a row you could not run gives the published figure with a citation and a footnote. This is ordinary scholarly practice, and it is the difference between a section that says nothing and a section that positions the run's contribution in the literature.
The rule you are working under is that no number in the report may be invented, estimated, or narrated into existence: every quantity traces to something. That rule is correct and this does not weaken it. A published value has a source; it is traceable; it is not an invention. What makes the difference is entirely in how it is carried:
synthesised candidate (ref. 4, Fig. 5)."
styled band or marker with the source in the legend — never a series that reads as one of yours. In a table, a source column.
the target your work is compared against, not evidence that your work is right.
A number carried that way is a citation. The same number carried as your own measurement is a fabrication. The line is bright and it is about labelling.
The near-miss is worth naming because it looks like the right answer. A run puts the published band on a figure, and then captions it "carried as an external reference, and never as a validation of this run" — and the section still reads as a refusal, because nothing of the run's own is placed against the band. A reader sees a number that belongs to someone else and no claim.
The version that works puts the run's own quantity in the same frame and states the distance: "our pipeline predicts 355 K against the 311-317 K they measured — 0.8σ of our uncertainty budget." Same published number, same honesty about who measured what, and now there is a result in the sentence.
The disclaimer is still correct and still belongs there. It goes after the comparison, as a qualification of it, not instead of it.
This is the part that turns a citation into a result. You could not synthesise the candidate — but you designed it, and you have a predicted property with an uncertainty. Put the prediction and the published measurement in the same figure and state the deviation. You could not run the licensed benchmark — but you have your metric on the open subset, and the source reports both, so the offset between them is estimable.
The section then reads: this is what was asked, this is what the field has measured, this is what this run predicts, this is the gap between them, and this is what would close it. That is a contribution. "Not attempted" is not.
In the section a reader looks in for the answer — under the heading the task's own words would send them to — not in Limitations. Limitations gets a cross-reference. A deliverable answered only in Limitations has been answered in the one place a reader goes to find out what the run failed at.
For each named deliverable the run could not produce: is there a section under a heading a reader would look for, containing the field's value with its source, this run's nearest evidence, and the distance between them? If the answer is "it is mentioned in Limitations", it is not done.
See also cover-what-the-task-named for enumerating deliverables in the first place, and citation-discipline for how the source is recorded once you cite it.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 54,927 | 21,703 | -60% | 1 | 1 | 0% | 3,096 | 4,241 | +37% | 0 | 0 | — |
case-02 | pass→fail | 17,832 | 21,108 | +18% | 1 | 1 | 0% | 2,779 | 4,191 | +51% | 0 | 0 | — |
case-03 | pass→pass | 12,408 | 11,966 | -4% | 1 | 1 | 0% | 1,821 | 2,731 | +50% | 0 | 0 | — |
case-04 | fail→fail | 31,473 | 29,013 | -8% | 1 | 1 | 0% | 2,117 | 3,721 | +76% | 0 | 0 | — |
case-05 | fail→pass | 24,791 | 25,094 | +1% | 1 | 1 | 0% | 2,801 | 4,473 | +60% | 0 | 0 | — |
case-06 | pass→pass | 12,633 | 14,441 | +14% | 1 | 1 | 0% | 1,776 | 2,987 | +68% | 0 | 0 | — |
case-07 | pass→pass | 16,217 | 13,017 | -20% | 1 | 1 | 0% | 1,904 | 3,013 | +58% | 0 | 0 | — |
case-08 | pass→pass | 22,664 | 13,253 | -42% | 1 | 1 | 0% | 2,410 | 3,069 | +27% | 0 | 0 | — |
case-09 | pass→pass | 17,910 | 15,960 | -11% | 1 | 1 | 0% | 2,146 | 3,110 | +45% | 0 | 0 | — |
case-10 | fail→fail | 15,327 | 14,325 | -7% | 1 | 1 | 0% | 2,235 | 3,213 | +44% | 0 | 0 | — |
case-11 | pass→pass | 33,593 | 40,056 | +19% | 1 | 1 | 0% | 2,045 | 3,189 | +56% | 0 | 0 | — |
case-12 | fail→pass | 39,047 | 14,602 | -63% | 1 | 1 | 0% | 1,653 | 3,052 | +85% | 0 | 0 | — |
case-13 | pass→pass | 13,780 | 35,133 | +155% | 1 | 1 | 0% | 2,086 | 3,462 | +66% | 0 | 0 | — |
case-14 | fail→pass | 10,795 | 10,873 | +1% | 1 | 1 | 0% | 1,545 | 2,744 | +78% | 0 | 0 | — |
case-15 | fail→pass | 23,901 | 14,532 | -39% | 1 | 1 | 0% | 2,481 | 3,049 | +23% | 0 | 0 | — |
case-16 | fail→pass | 14,211 | 11,266 | -21% | 1 | 1 | 0% | 2,136 | 2,683 | +26% | 0 | 0 | — |
case-17 | fail→pass | 13,097 | 10,701 | -18% | 1 | 1 | 0% | 1,918 | 2,646 | +38% | 0 | 0 | — |
case-18 | fail→pass | 17,486 | 13,133 | -25% | 1 | 1 | 0% | 2,398 | 2,888 | +20% | 0 | 0 | — |
case-19 | fail→pass | 12,992 | 13,327 | +3% | 1 | 1 | 0% | 1,819 | 3,077 | +69% | 0 | 0 | — |
case-20 | fail→pass | 31,882 | 18,082 | -43% | 1 | 1 | 0% | 2,234 | 3,693 | +65% | 0 | 0 | — |
case-21 | fail→pass | 14,860 | 11,465 | -23% | 1 | 1 | 0% | 2,080 | 2,574 | +24% | 0 | 0 | — |
case-22 | pass→pass | 11,000 | 23,108 | +110% | 1 | 1 | 0% | 1,733 | 4,770 | +175% | 0 | 0 | — |
case-23 | pass→pass | 15,965 | 17,650 | +11% | 1 | 1 | 0% | 1,542 | 3,422 | +122% | 0 | 0 | — |
case-24 | pass→pass | 12,833 | 17,972 | +40% | 1 | 1 | 0% | 2,014 | 3,578 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +38 percentage points is the difference between those two pass rates over the 24 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.