Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis when an energy-system result rests on an aggregation, a scenario or a saving. Covers the counterfactual pair a saving has to be quoted against, checking that a hierarchy sums, and reporting at the data's native temporal resolution.
.claude/skills/tangxiangru-energy-counterfactual-pair-and-hierarchy-closure/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -1% | 0% |
Energy-systems work is graded on accounting. Three habits outsiders skip.
The counterfactual pair. No headline number stands alone: solve or measure the system twice, once with the studied mechanism active and once with it removed (network constraints relaxed, policy off, technology absent, do-nothing ideal), under identical cost and physical accounting, and report both arms and their difference. Design the counterfactual run at the same time as the main run; it is half the result, not a sensitivity check.
Nested closure. Report at every level the system defines - asset, then building, zone, community or country, then whole system - and show the closure: components must sum to the reported total, with the residual in physical units and an overlay of the two series for one representative window. State every rate or fraction with its numerator, its denominator and both in physical units, naming the population the denominator covers: which assets, which hours, which sites.
Native resolution. Plot the full horizon at the data's own timestep and the full extent at its own spatial unit, decomposing the total into its components at each timestep, rather than reporting only period aggregates or one mean map.
Then report one quantified effect for every input variable and data layer you were supplied, including those that turn out not to matter: an explicit null is a result, a variable that vanishes from the report reads as an omitted mechanism. Report the full cross-product of scenarios and name the best and worst cell, not the mean.
Four of four Energy tasks demand a paired counterfactual and three demand hierarchical closure - and both were the specific misses. The constrained-vs-unconstrained contrast existed only as two rows inside a wide scenario table and the judge wrote 'there is no explicit unconstrained case'; the highest-weight criterion of the worst task (0.40) was the children-sum-to-parent identity, computed to 1e-16 but used as one clause of a data-audit argument with no section, figure or day-level overlay, scoring 0. The per-layer null-effect rule recovers the criterion lost when a supplied input layer was discarded and then reported as a data limitation.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 48,984 | 42,228 | -14% | 1 | 1 | 0% | 8,290 | 8,762 | +6% | 0 | 0 | — |
case-02 | fail→fail | 47,194 | 62,211 | +32% | 1 | 1 | 0% | 8,287 | 8,759 | +6% | 0 | 0 | — |
case-03 | fail→fail | 53,283 | 46,746 | -12% | 1 | 1 | 0% | 8,427 | 8,754 | +4% | 0 | 0 | — |
case-04 | pass→pass | 11,328 | 27,944 | +147% | 1 | 1 | 0% | 1,671 | 5,311 | +218% | 0 | 0 | — |
case-05 | pass→pass | 14,577 | 38,885 | +167% | 1 | 1 | 0% | 2,532 | 7,784 | +207% | 0 | 0 | — |
case-06 | pass→fail | 32,689 | 60,864 | +86% | 1 | 1 | 0% | 5,102 | 10,726 | +110% | 0 | 0 | — |
case-07 | fail→fail | 33,032 | 65,194 | +97% | 1 | 1 | 0% | 5,710 | 12,426 | +118% | 0 | 0 | — |
case-08 | fail→pass | 38,607 | 91,538 | +137% | 1 | 1 | 0% | 7,416 | 8,720 | +18% | 0 | 0 | — |
case-09 | fail→pass | 36,301 | 49,488 | +36% | 1 | 1 | 0% | 5,601 | 8,730 | +56% | 0 | 0 | — |
case-10 | fail→pass | 28,493 | 47,510 | +67% | 1 | 1 | 0% | 4,036 | 8,737 | +116% | 0 | 0 | — |
case-11 | fail→fail | 58,120 | 76,139 | +31% | 1 | 1 | 0% | 9,737 | 8,729 | -10% | 0 | 0 | — |
case-12 | fail→pass | 105,864 | 46,445 | -56% | 1 | 1 | 0% | 9,776 | 8,736 | -11% | 0 | 0 | — |
case-13 | fail→fail | 48,988 | 58,254 | +19% | 1 | 1 | 0% | 8,245 | 10,042 | +22% | 0 | 0 | — |
case-14 | fail→pass | 54,045 | 44,913 | -17% | 1 | 1 | 0% | 8,792 | 8,726 | -1% | 0 | 0 | — |
case-15 | fail→fail | 59,876 | 128,224 | +114% | 1 | 1 | 0% | 10,589 | 10,095 | -5% | 0 | 0 | — |
case-16 | fail→fail | 47,071 | 44,128 | -6% | 1 | 1 | 0% | 7,876 | 8,721 | +11% | 0 | 0 | — |
case-17 | fail→pass | 38,712 | 64,932 | +68% | 1 | 1 | 0% | 5,904 | 11,775 | +99% | 0 | 0 | — |
case-18 | fail→fail | 46,068 | 48,275 | +5% | 1 | 1 | 0% | 8,240 | 8,712 | +6% | 0 | 0 | — |
case-19 | fail→fail | 18,527 | 43,973 | +137% | 1 | 1 | 0% | 3,100 | 8,716 | +181% | 0 | 0 | — |
case-20 | fail→pass | 52,359 | 47,853 | -9% | 1 | 1 | 0% | 8,593 | 8,519 | -1% | 0 | 0 | — |
case-21 | fail→fail | 17,835 | 78,431 | +340% | 1 | 1 | 0% | 2,802 | 8,711 | +211% | 0 | 0 | — |
case-22 | fail→fail | 38,386 | 41,735 | +9% | 1 | 1 | 0% | 6,711 | 8,729 | +30% | 0 | 0 | — |
case-23 | fail→fail | 43,037 | 48,814 | +13% | 1 | 1 | 0% | 8,265 | 9,688 | +17% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.