Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis and writing when a result rests on a fit or a calibration chain and you are about to quote it with a single uncertainty. Covers itemising the error budget term by term, keeping the fit's own bookkeeping visible, and why the audit trail is the result a referee checks first.
.claude/skills/tangxiangru-astronomy-error-budget-is-the-audit-trail/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 192% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 168% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 102% | 0% |
Decide the uncertainty machinery at design time, because it dictates what you must store. Propagate from the full measurement distribution, Monte Carlo over posterior samples or a full covariance matrix, never from a point estimate plus an error bar, and name the propagation method and which correlations it carries and which it drops.
Where a fit is involved, state its bookkeeping: how many constraint equations, broken down by category; how many free parameters; the resulting degrees of freedom; chi-square per degree of freedom; and whether the covariance used was diagonal or full. Read the reduced chi-square as a statement about whether the error budget matches the residuals. Then draw the residual diagnostic: model minus each individual measurement, one point per constraint, with its error bar and a zero line. Astronomers read that panel first, and an rms quoted in prose does not replace it.
Characterise the input catalogue quantitatively before modelling: sample count, central value and dispersion, correlations between inputs, units, and the assumed auxiliary quantities nobody measured. Report mean with standard deviation and median with percentile interval; subfields differ on which is conventional.
State the headline as value plus or minus uncertainty with units, an explicit confidence level, the relative precision, the object or range it is conditioned on, and the discrepancy from the community value in sigma. If you conclude the inputs cannot support the published value, still emit that value, its diagnostic figure and the structural counts in this form, and report the discrepancy beside them.
Astronomy's mechanism 3 (task reframing) cost all three criteria of one task, 25-28 each, because the target-form deliverables (headline value at the stated precision, the equation/parameter/dof accounting, the residual diagnostic) were never emitted while the agent wrote a forensic audit; the final clause forces them out anyway. Mechanism 4 (convention mismatch: median with credible interval where mean plus sd is the field's habit) cost about 15 points on an otherwise strong criterion, fixed by reporting both. Mechanism 1's skipped cheap first step, characterising the input layer in standard form, is the opening clause.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 39,566 | 88,124 | +123% | 1 | 1 | 0% | 8,298 | 13,834 | +67% | 0 | 0 | — |
case-02 | fail→pass | 38,788 | 70,640 | +82% | 1 | 1 | 0% | 7,051 | 6,846 | -3% | 0 | 0 | — |
case-03 | fail→fail | 91,970 | 53,453 | -42% | 1 | 1 | 0% | 8,293 | 9,782 | +18% | 0 | 0 | — |
case-04 | fail→pass | 10,941 | 24,496 | +124% | 1 | 1 | 0% | 1,709 | 4,983 | +192% | 0 | 0 | — |
case-05 | fail→pass | 41,259 | 48,890 | +18% | 1 | 1 | 0% | 2,152 | 5,767 | +168% | 0 | 0 | — |
case-06 | fail→fail | 28,520 | 30,434 | +7% | 1 | 1 | 0% | 3,666 | 6,438 | +76% | 0 | 0 | — |
case-07 | fail→pass | 19,076 | 33,594 | +76% | 1 | 1 | 0% | 3,223 | 6,523 | +102% | 0 | 0 | — |
case-08 | pass→pass | 19,537 | 49,730 | +155% | 1 | 1 | 0% | 3,334 | 10,020 | +201% | 0 | 0 | — |
case-09 | fail→pass | 5,916 | 37,120 | +527% | 1 | 1 | 0% | 1,046 | 5,317 | +408% | 0 | 0 | — |
case-10 | fail→pass | 15,480 | 26,160 | +69% | 1 | 1 | 0% | 2,355 | 5,188 | +120% | 0 | 0 | — |
case-11 | fail→pass | 4,719 | 43,796 | +828% | 1 | 1 | 0% | 528 | 6,296 | +1092% | 0 | 0 | — |
case-12 | fail→pass | 3,926 | 34,940 | +790% | 1 | 1 | 0% | 390 | 7,332 | +1780% | 0 | 0 | — |
case-13 | fail→pass | 7,178 | 55,498 | +673% | 1 | 1 | 0% | 1,056 | 4,863 | +361% | 0 | 0 | — |
case-14 | pass→pass | 14,180 | 47,251 | +233% | 1 | 1 | 0% | 2,155 | 5,387 | +150% | 0 | 0 | — |
case-15 | fail→fail | 22,359 | 22,455 | +0% | 1 | 1 | 0% | 2,104 | 4,344 | +106% | 0 | 0 | — |
case-16 | pass→pass | 14,607 | 47,572 | +226% | 1 | 1 | 0% | 2,536 | 5,145 | +103% | 0 | 0 | — |
case-17 | pass→pass | 34,177 | 59,987 | +76% | 1 | 1 | 0% | 2,269 | 7,510 | +231% | 0 | 0 | — |
case-18 | pass→pass | 14,709 | 19,981 | +36% | 1 | 1 | 0% | 2,393 | 3,921 | +64% | 0 | 0 | — |
case-19 | pass→pass | 26,112 | 47,401 | +82% | 1 | 1 | 0% | 2,942 | 7,974 | +171% | 0 | 0 | — |
case-20 | fail→pass | 13,565 | 64,807 | +378% | 1 | 1 | 0% | 2,045 | 8,484 | +315% | 0 | 0 | — |
case-21 | fail→pass | 48,250 | 36,494 | -24% | 1 | 1 | 0% | 1,174 | 7,223 | +515% | 0 | 0 | — |
case-22 | pass→pass | 13,953 | 31,296 | +124% | 1 | 1 | 0% | 1,957 | 5,791 | +196% | 0 | 0 | — |
case-23 | pass→fail | 33,849 | 48,072 | +42% | 1 | 1 | 0% | 7,297 | 8,722 | +20% | 0 | 0 | — |
case-24 | pass→fail | 12,461 | 20,420 | +64% | 1 | 1 | 0% | 1,802 | 4,065 | +126% | 0 | 0 | — |
case-25 | pass→fail | 39,479 | 72,496 | +84% | 1 | 1 | 0% | 3,798 | 9,601 | +153% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 24 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.