Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design when you are about to run an improved variant of a system before its default configuration, or fold several claims into one panel. Covers running the default recipe with its conventional diagnostics first, and giving each claim its own plain panel.
.claude/skills/tangxiangru-energy-canonical-configuration-before-the-enhanced-variant/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 37% | 0% |
Two things get engineering and energy reproductions marked incomplete.
The default configuration. Implement the field's canonical recipe exactly as specified - the given sample count, architecture, optimiser, budget, convergence tolerance, solver settings - and report its conventional diagnostics: the loss actually optimised with its start and end values, iteration or epoch count, how many model evaluations were attempted and how many converged, wall time per pipeline stage, and the speed-up claimed against the method it replaces. Only then run your improved variant, and present it as an ablation against that baseline. A study that goes straight to the better design leaves every diagnostic the field expects unreported, however superior it is. Use the field's published error definitions rather than substituting your preferred statistic.
A named method may not enter a cut order, and a design is not frozen until it has run. Pilot the box with 24-48 real evaluations before committing to it, and record attempted, succeeded and the failure mode. A 37.5% pilot yield is not a caveat to carry forward, it is the design telling you it is the wrong box: one run recorded exactly that, ruled it "a guard and not a re-design", froze the preregistration, and shipped 27.1% convergence to the reader.
Error per unit, not only pooled: a true-versus-estimated table with one row per estimated parameter, per site, per asset or per variable, carrying an absolute and a relative-error column, alongside the aggregate residual metric.
Figures. Give every headline claim its own single-purpose panel in the canonical diagnostic form: a two-series overlay when claiming agreement, a stacked decomposition when claiming a total splits, a utilisation-ratio time series when claiming a limit binds, a map at the native spatial unit when claiming heterogeneity. A dense multi-panel composite of novel diagnostics is an addition, never the only place a claim appears. And if you judge the supplied inputs unrepresentative, still run the pre-specified analysis on them and report it first, with the audit in a clearly separate later section.
On the inverse-identification task the agent ran a far better study than the reference - large design of experiments, deep ensemble, Sobol screen - and never ran the textbook small-design plain-MLP configuration, so every criterion asking for the standard recipe's diagnostics (training MSE trajectory, epoch count, run yield, per-parameter true-vs-identified error) scored 0 or 25; it reported held-out R-squared where the field reports training MSE and per-parameter percent error. Separately, three of the four zero-scored Energy criteria were single plain panels that no 4-panel composite contained. The audit-ordering clause addresses the run that discarded the supplied files entirely and scored 8.4.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 52,453 | 41,840 | -20% | 1 | 1 | 0% | 9,513 | 7,241 | -24% | 0 | 0 | — |
case-02 | fail→pass | 67,081 | 45,052 | -33% | 1 | 1 | 0% | 10,614 | 7,615 | -28% | 0 | 0 | — |
case-03 | fail→fail | 42,851 | 43,245 | +1% | 1 | 1 | 0% | 8,271 | 8,881 | +7% | 0 | 0 | — |
case-04 | fail→pass | 36,481 | 24,179 | -34% | 1 | 1 | 0% | 2,790 | 4,039 | +45% | 0 | 0 | — |
case-05 | fail→pass | 44,428 | 33,349 | -25% | 1 | 1 | 0% | 1,571 | 4,427 | +182% | 0 | 0 | — |
case-06 | pass→pass | 17,773 | 15,019 | -15% | 1 | 1 | 0% | 2,533 | 2,742 | +8% | 0 | 0 | — |
case-07 | pass→pass | 22,155 | 19,401 | -12% | 1 | 1 | 0% | 3,561 | 3,863 | +8% | 0 | 0 | — |
case-08 | pass→pass | 21,068 | 11,839 | -44% | 1 | 1 | 0% | 3,090 | 2,439 | -21% | 0 | 0 | — |
case-09 | pass→pass | 19,309 | 11,986 | -38% | 1 | 1 | 0% | 2,971 | 2,508 | -16% | 0 | 0 | — |
case-10 | pass→pass | 18,073 | 10,531 | -42% | 1 | 1 | 0% | 2,411 | 1,933 | -20% | 0 | 0 | — |
case-11 | pass→pass | 29,362 | 10,824 | -63% | 1 | 1 | 0% | 2,950 | 2,002 | -32% | 0 | 0 | — |
case-12 | fail→pass | 18,963 | 18,698 | -1% | 1 | 1 | 0% | 2,410 | 3,298 | +37% | 0 | 0 | — |
case-13 | fail→fail | 16,163 | 13,381 | -17% | 1 | 1 | 0% | 2,414 | 2,509 | +4% | 0 | 0 | — |
case-14 | fail→pass | 16,112 | 12,368 | -23% | 1 | 1 | 0% | 2,139 | 2,384 | +11% | 0 | 0 | — |
case-15 | pass→pass | 18,712 | 8,447 | -55% | 1 | 1 | 0% | 2,817 | 1,875 | -33% | 0 | 0 | — |
case-16 | pass→pass | 17,147 | 16,473 | -4% | 1 | 1 | 0% | 2,438 | 3,160 | +30% | 0 | 0 | — |
case-17 | fail→pass | 21,910 | 21,800 | -1% | 1 | 1 | 0% | 2,722 | 2,972 | +9% | 0 | 0 | — |
case-18 | fail→pass | 20,619 | 18,388 | -11% | 1 | 1 | 0% | 3,014 | 3,044 | +1% | 0 | 0 | — |
case-19 | fail→pass | 19,271 | 10,587 | -45% | 1 | 1 | 0% | 2,992 | 2,171 | -27% | 0 | 0 | — |
case-20 | pass→pass | 38,946 | 42,647 | +10% | 1 | 1 | 0% | 8,233 | 8,843 | +7% | 0 | 0 | — |
case-21 | pass→pass | 14,949 | 31,336 | +110% | 1 | 1 | 0% | 2,858 | 7,016 | +145% | 0 | 0 | — |
case-22 | pass→pass | 19,411 | 43,200 | +123% | 1 | 1 | 0% | 3,821 | 8,869 | +132% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.