Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and analysis when a physical quantity is about to be reported from one estimator, or an uncertainty quoted without propagation. Covers measuring it a second independent way, propagating the error through the chain, and generating the observable forward from the fitted model to check it.
.claude/skills/tangxiangru-physics-two-estimators-propagation-and-a-forward-model/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 212% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 56% | 0% |
Matching a published measurement is the floor in physics; you exceed it in four specific ways, so design for them up front.
Measure the same quantity through two or more independent estimators or measurement channels and quantify their agreement with a paired plot, residuals and a chi-square using per-point uncertainties, rather than asserting it. Where one channel reaches beyond the other, validate the proxy on the overlap region before relying on it outside.
Attach an uncertainty to every number and a propagated band to every model curve drawn against data, obtained by Monte Carlo or analytic propagation from the inputs' own uncertainties. For ensemble points, state how many members contributed and confirm none were dropped.
Generate data rather than only re-fitting the observation: at least one result should come from a stated forward model, an energy minimisation, a kinetic simulation of the process, an error-budget model, compared against the measurement. When the claim concerns a process, plot the process itself, observable against time, step or index with the state label on it, plus the event-type statistics; final-state energies do not evidence a mechanism.
Re-derive derived quantities from the rawest artifact available, treating shipped peak positions, landmark tables and summary files as claims to verify. Check absolute scale and units against known physical bounds and an independently computed characteristic scale. Sweep each control variable separately with the others held fixed, both directions where the data form a grid. Give every claim one self-describing figure and put its number in the summary.
Eight of Physics's fourteen items sit in the 45-60 parity band, and moving that cluster is worth about +10 discipline points; the measured route above parity is more independent estimators, tighter propagated uncertainties, wider sweeps and validation of the estimator's own assumptions, none of which the bare agent did systematically. The process-figure clause targets the one near-absent item in the discipline (score 5, weight 0.30): the agent ran the simulation and reproduced the effect but plotted final-state energies instead of the process, and never computed the event-type statistics.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 36,592 | 46,342 | +27% | 1 | 1 | 0% | 6,204 | 8,704 | +40% | 0 | 0 | — |
case-02 | fail→pass | 18,886 | 46,208 | +145% | 1 | 1 | 0% | 2,790 | 8,705 | +212% | 0 | 0 | — |
case-03 | fail→fail | 88,525 | 53,158 | -40% | 1 | 1 | 0% | 8,418 | 9,946 | +18% | 0 | 0 | — |
case-04 | fail→pass | 18,660 | 57,690 | +209% | 1 | 1 | 0% | 2,743 | 5,961 | +117% | 0 | 0 | — |
case-05 | pass→fail | 16,333 | 16,507 | +1% | 1 | 1 | 0% | 2,193 | 2,488 | +13% | 0 | 0 | — |
case-06 | fail→fail | 15,204 | 19,140 | +26% | 1 | 1 | 0% | 2,103 | 2,980 | +42% | 0 | 0 | — |
case-07 | fail→pass | 16,160 | 15,495 | -4% | 1 | 1 | 0% | 2,309 | 2,750 | +19% | 0 | 0 | — |
case-08 | fail→pass | 16,510 | 19,566 | +19% | 1 | 1 | 0% | 2,496 | 3,431 | +37% | 0 | 0 | — |
case-09 | fail→pass | 16,233 | 22,129 | +36% | 1 | 1 | 0% | 2,236 | 3,498 | +56% | 0 | 0 | — |
case-10 | fail→pass | 18,548 | 28,068 | +51% | 1 | 1 | 0% | 2,487 | 4,556 | +83% | 0 | 0 | — |
case-11 | fail→pass | 13,967 | 14,723 | +5% | 1 | 1 | 0% | 1,994 | 2,529 | +27% | 0 | 0 | — |
case-12 | pass→pass | 18,642 | 33,038 | +77% | 1 | 1 | 0% | 2,969 | 5,200 | +75% | 0 | 0 | — |
case-13 | fail→pass | 21,958 | 22,299 | +2% | 1 | 1 | 0% | 3,061 | 3,893 | +27% | 0 | 0 | — |
case-14 | fail→fail | 39,913 | 41,553 | +4% | 1 | 1 | 0% | 8,219 | 8,670 | +5% | 0 | 0 | — |
case-15 | pass→pass | 21,355 | 44,372 | +108% | 1 | 1 | 0% | 3,733 | 8,272 | +122% | 0 | 0 | — |
case-16 | pass→pass | 25,602 | 45,953 | +79% | 1 | 1 | 0% | 3,827 | 8,680 | +127% | 0 | 0 | — |
case-17 | pass→pass | 18,739 | 45,220 | +141% | 1 | 1 | 0% | 3,101 | 8,672 | +180% | 0 | 0 | — |
case-18 | fail→pass | 17,269 | 30,000 | +74% | 1 | 1 | 0% | 2,762 | 5,002 | +81% | 0 | 0 | — |
case-19 | pass→pass | 17,651 | 25,214 | +43% | 1 | 1 | 0% | 2,813 | 5,182 | +84% | 0 | 0 | — |
case-20 | fail→pass | 22,527 | 39,408 | +75% | 1 | 1 | 0% | 3,299 | 6,193 | +88% | 0 | 0 | — |
case-21 | fail→pass | 17,018 | 19,238 | +13% | 1 | 1 | 0% | 2,549 | 3,162 | +24% | 0 | 0 | — |
case-22 | fail→pass | 16,506 | 51,155 | +210% | 1 | 1 | 0% | 2,413 | 8,997 | +273% | 0 | 0 | — |
case-23 | fail→fail | 20,192 | 26,518 | +31% | 1 | 1 | 0% | 2,848 | 4,623 | +62% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +48 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.