Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis when a fit is about to be reported as the answer without a rival model being excluded. Covers naming the competing model families, showing which the data rules out, and treating the fit protocol — range, weighting, priors — as part of the result rather than as a setting.
.claude/skills/tangxiangru-physics-discriminate-model-families-and-defend-the-fit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-01 | ✓→✗ | ▼ Worse | -85% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 172% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 80% | 0% |
| case-23 | ✓→✓ | = Same ✓ | -3% | 0% |
Name the competing model families before you touch the data. A physics result is not "our number is good"; it is "quantity Q, measured against control X with everything else fixed and carrying an uncertainty, follows family A and excludes family B". Fit every candidate family to the same data on a common footing, report a per-family goodness-of-fit, and state which families are excluded and at what confidence. The supplied materials usually name the alternatives; where they do not, the conventional theory and a trivial null are still required arms.
When the reported number is a parameter of a functional form, an exponent, a slope, a critical value, treat the fit protocol as part of the result. That value depends on the abscissa, the fit window, the weighting, and which reference quantities (amplitude, critical point, offset) are held fixed rather than free. Enumerate those choices, report the parameter under the field's standard convention first, and attach a sensitivity table over the alternatives. Test the fixed-exponent model with a goodness-of-fit statistic in addition to free-fitting the exponent: the two answer different questions and can disagree by a factor of two when the small-signal end is noise-dominated.
Run the original protocol on its own terms and report that outcome as the headline for each target result. Data-provenance problems and alternative interpretations belong in a subordinate section, and "the effect is present with the wrong functional form" must be stated separately from "the effect is absent".
Physics has 0% absent, so its headroom is the five items scoring 28-38, every one of them a case where the agent's number contradicted the paper through fit-protocol choice rather than through error. In one task the criterion's own fixed-form goodness-of-fit test passes on the shipped data (R-squared 0.9886) while the agent free-fitted a log-log exponent, got a different value, and declared the effect refuted; running both would have kept the item. The model-family clause targets the discipline's constant claim grammar, which three of four tasks state explicitly as two or more named alternatives.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 44,017 | 12,794 | -71% | 1 | 1 | 0% | 8,348 | 1,225 | -85% | 0 | 0 | — |
case-02 | fail→fail | 57,611 | 6,821 | -88% | 1 | 1 | 0% | 8,335 | 731 | -91% | 0 | 0 | — |
case-03 | fail→fail | 12,979 | 11,983 | -8% | 1 | 1 | 0% | 342 | 1,152 | +237% | 0 | 0 | — |
case-04 | fail→fail | 42,347 | 16,033 | -62% | 1 | 1 | 0% | 326 | 1,023 | +214% | 0 | 0 | — |
case-05 | fail→fail | 24,937 | 18,956 | -24% | 1 | 1 | 0% | 333 | 1,054 | +217% | 0 | 0 | — |
case-06 | fail→fail | 18,629 | 7,966 | -57% | 1 | 1 | 0% | 3,490 | 960 | -72% | 0 | 0 | — |
case-07 | fail→fail | 18,687 | 8,858 | -53% | 1 | 1 | 0% | 3,094 | 631 | -80% | 0 | 0 | — |
case-08 | fail→fail | 31,257 | 8,809 | -72% | 1 | 1 | 0% | 5,509 | 787 | -86% | 0 | 0 | — |
case-09 | fail→pass | 27,685 | 40,766 | +47% | 1 | 1 | 0% | 4,410 | 7,850 | +78% | 0 | 0 | — |
case-10 | fail→fail | 50,365 | 9,287 | -82% | 1 | 1 | 0% | 3,267 | 1,057 | -68% | 0 | 0 | — |
case-11 | fail→fail | 76,418 | 13,243 | -83% | 1 | 1 | 0% | 3,994 | 904 | -77% | 0 | 0 | — |
case-12 | fail→fail | 60,259 | 13,576 | -77% | 1 | 1 | 0% | 4,077 | 1,026 | -75% | 0 | 0 | — |
case-13 | fail→fail | 10,264 | 12,811 | +25% | 1 | 1 | 0% | 510 | 1,445 | +183% | 0 | 0 | — |
case-14 | fail→fail | 23,571 | 13,151 | -44% | 1 | 1 | 0% | 3,969 | 816 | -79% | 0 | 0 | — |
case-15 | fail→fail | 36,560 | 13,558 | -63% | 1 | 1 | 0% | 3,311 | 1,068 | -68% | 0 | 0 | — |
case-16 | fail→fail | 73,597 | 8,590 | -88% | 1 | 1 | 0% | 3,925 | 943 | -76% | 0 | 0 | — |
case-17 | fail→fail | 19,093 | 10,517 | -45% | 1 | 1 | 0% | 3,291 | 881 | -73% | 0 | 0 | — |
case-18 | fail→fail | 27,576 | 10,999 | -60% | 1 | 1 | 0% | 3,189 | 770 | -76% | 0 | 0 | — |
case-19 | fail→fail | 17,405 | 16,112 | -7% | 1 | 1 | 0% | 2,903 | 1,082 | -63% | 0 | 0 | — |
case-20 | fail→fail | 64,136 | 12,797 | -80% | 1 | 1 | 0% | 5,290 | 1,008 | -81% | 0 | 0 | — |
case-21 | pass→fail | 32,551 | 43,898 | +35% | 1 | 1 | 0% | 3,212 | 8,736 | +172% | 0 | 0 | — |
case-22 | pass→fail | 21,137 | 40,094 | +90% | 1 | 1 | 0% | 4,856 | 8,752 | +80% | 0 | 0 | — |
case-23 | pass→pass | 22,759 | 20,237 | -11% | 1 | 1 | 0% | 3,505 | 3,392 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 4 counted toward the lift figure. The other 19 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -33 percentage points is the difference between those two pass rates over the 4 comparable cases. 7 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.