Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and implementation when a protocol is specified and you have found a reason to deviate, or when a pipeline stage is about to run without its conventional diagnostic. Covers running the protocol as specified as the foreground result, and leaving every stage's default panel behind you.
.claude/skills/tangxiangru-material-as-specified-run-and-stage-diagnostics/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 152% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -7% | 0% |
Materials pipelines are graded stage by stage, so run the protocol exactly as specified before you improve it: the given parameter values, cell sizes, sample counts, budgets and convergence thresholds. That run is the foreground result, reported in the protocol's own units. Nearly every supplied specification has a defect you will find; the audit belongs in a clearly separated second analysis with the delta attributed, and must not take the title, the abstract's first sentence, or panel (a). A lead figure captioned with what is wrong with the spec is read as a declined reproduction even when the correct number sits in a table two sections later.
Emit each stage's default diagnostic panel even when a deeper analysis supersedes it: input characterisation (N per split, class balance, target range, replicate noise), training and validation objective versus epoch, held-out metric versus step, parity plot against the reference with the identity line, best-so-far versus evaluation count, generated-versus-reference scatter in the physical parameter space. These panels are cheap, conventional, and their absence is unrecoverable.
A named method may not enter a cut order. When the budget will not carry everything, cut the extension, the extra seeds, the second substrate — and run the paper's own method at one seed and fewer epochs instead. A reproduction at a fifth of the scale scores; a reproduction replaced by a cheaper engine does not. One run listed the paper's graph encoder as item (5) of its cut order, took the cut, shipped gradient boosting as the headline, and watched that criterion go 28 -> 5 against a plain agent's 38.
When the specification names a method family -- a particular surrogate, optimiser, simulator or architecture -- run that one and report it, then place your alternative beside it rather than instead of it. When the final validation needs an instrument you lack (synthesis, a measurement, a high-fidelity simulation), do not drop the stage: report the best available proxy, label it a proxy, quantify its uncertainty, and give it its own panel.
Targets the four Material mechanisms that produce 0-35 scores with no absence: task substitution (all four bare runs detected a defective spec and made the audit the headline, so the demanded number existed nowhere), framing inversion (the correct result relegated behind a 'the spec is broken' lead panel scored 25), the never-plotted standard diagnostic (a weight-0.5 criterion scored 5 and 0 because loss- and accuracy-versus-epoch traces were judged too boring to draw), and the silently dropped wet-lab/expensive-simulation stage.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 89,499 | 48,315 | -46% | 1 | 1 | 0% | 8,319 | 9,478 | +14% | 0 | 0 | — |
case-02 | pass→pass | 32,975 | 37,050 | +12% | 1 | 1 | 0% | 5,177 | 6,888 | +33% | 0 | 0 | — |
case-03 | fail→fail | 48,335 | 47,257 | -2% | 1 | 1 | 0% | 8,946 | 6,283 | -30% | 0 | 0 | — |
case-04 | fail→pass | 11,340 | 16,793 | +48% | 1 | 1 | 0% | 1,674 | 3,417 | +104% | 0 | 0 | — |
case-05 | fail→fail | 22,180 | 14,293 | -36% | 1 | 1 | 0% | 1,885 | 2,537 | +35% | 0 | 0 | — |
case-06 | fail→pass | 46,775 | 21,530 | -54% | 1 | 1 | 0% | 1,051 | 2,652 | +152% | 0 | 0 | — |
case-07 | fail→pass | 57,053 | 10,500 | -82% | 1 | 1 | 0% | 2,415 | 2,213 | -8% | 0 | 0 | — |
case-08 | fail→pass | 14,359 | 10,969 | -24% | 1 | 1 | 0% | 1,983 | 1,840 | -7% | 0 | 0 | — |
case-09 | pass→pass | 24,477 | 10,134 | -59% | 1 | 1 | 0% | 2,122 | 2,005 | -6% | 0 | 0 | — |
case-10 | fail→fail | 11,953 | 20,806 | +74% | 1 | 1 | 0% | 1,650 | 1,997 | +21% | 0 | 0 | — |
case-11 | fail→pass | 12,682 | 10,305 | -19% | 1 | 1 | 0% | 1,665 | 2,095 | +26% | 0 | 0 | — |
case-12 | fail→fail | 29,816 | 27,690 | -7% | 1 | 1 | 0% | 2,446 | 2,160 | -12% | 0 | 0 | — |
case-13 | pass→pass | 35,904 | 13,246 | -63% | 1 | 1 | 0% | 2,040 | 1,520 | -25% | 0 | 0 | — |
case-14 | pass→pass | 31,205 | 16,821 | -46% | 1 | 1 | 0% | 2,752 | 3,208 | +17% | 0 | 0 | — |
case-15 | fail→pass | 12,156 | 10,557 | -13% | 1 | 1 | 0% | 1,609 | 2,111 | +31% | 0 | 0 | — |
case-16 | fail→fail | 27,132 | 32,644 | +20% | 1 | 1 | 0% | 2,478 | 2,675 | +8% | 0 | 0 | — |
case-17 | fail→pass | 15,382 | 9,376 | -39% | 1 | 1 | 0% | 2,231 | 1,931 | -13% | 0 | 0 | — |
case-18 | fail→pass | 19,087 | 24,227 | +27% | 1 | 1 | 0% | 1,932 | 2,577 | +33% | 0 | 0 | — |
case-19 | pass→pass | 14,694 | 15,497 | +5% | 1 | 1 | 0% | 1,850 | 1,452 | -22% | 0 | 0 | — |
case-20 | pass→pass | 18,059 | 22,312 | +24% | 1 | 1 | 0% | 3,134 | 4,434 | +41% | 0 | 0 | — |
case-21 | fail→fail | 6,692 | 9,223 | +38% | 1 | 1 | 0% | 861 | 1,883 | +119% | 0 | 0 | — |
case-22 | pass→pass | 54,308 | 59,908 | +10% | 1 | 1 | 0% | 3,468 | 5,245 | +51% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.