Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at Stage 06 and again before the report is finalised, when deciding which of the run's results enter the deliverable. Sweeps the run's own outputs for quantities it computed and never published, and covers the three shapes that sweep finds — the diagnostic never persisted, the column requested and dropped, the feasibility measurement discarded — and what to promote out of an appendix.
.claude/skills/tangxiangru-publish-what-the-run-already-computed/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -5% | 0% |
The work was done. The compute was spent. The number is on disk. And the report does not contain it — because the report's contents were chosen by the run's own hypothesis structure, and a quantity produced in service of a hypothesis but not adjudicating one has no section to live in.
This skill is the sweep that finds those. It is about the run's outputs; the-supplied-item-is-the-graded-unit is about the inputs — the named objects in data/ — and covers printing those objects rather than pointing at them. Run both. This one catches what that one cannot: a quantity nothing shipped and nothing asked for by name, which the run computed anyway because the analysis needed it.
Before you fix a section order or polish a sentence, do this once:
what it holds — not the path, the content: "per-epoch pre-training loss", "sequence identity per aligned chain pair", "the 497-row permutation-importance table", "seconds per optimiser step for the long-range and short-range heads".
report for a number out of the file; if none is there, the entry is empty.
the source study's own results, look for this? If yes, it goes in the body before you do anything else. If no, it may stay unpublished.
The sweep takes minutes and it is the highest-yield thing available at this point in the run, because everything it finds is already paid for.
The diagnostic that was never persisted. A training routine accumulates loss per epoch, returns it, and writes it to no artifact — so the convergence curve the source study made its case with is absent from a run that had the numbers in memory. Any quantity the source plots is worth persisting the moment it is computed, even when your own argument does not need it. This one is the most expensive because it is unrecoverable by Stage 07: there is nothing on disk to promote.
The column that was requested and dropped. A tool is asked for an extra field, the field lands in a CSV, and the report quotes everything from that CSV except that column. If some part of the run thought to ask for it, some part of the run thought it mattered.
The measurement taken for a feasibility check. Cost, timing and scaling numbers get produced to decide whether something is affordable, then discarded once the decision is made — on a report that trained a hundred models and says nothing about what any of them cost.
An object tabulated in an appendix and never analysed is half-published. Where a table has structure worth a statistic — nine superposition vectors have a dispersion, a set of per-pair scores has a distribution — computing it is a handful of lines and turns a dump into a result.
Not a licence to publish everything. Bulk arrays, intermediate caches and per-row dumps belong exactly where they are. The test is whether a reader would look for it: an object the task named, or a quantity the source study reports, belongs in the body; everything else is judged on whether it carries an argument.
See also cover-what-the-task-named for building the list this sweep checks against, the-supplied-item-is-the-graded-unit for the shipped objects, and result-table for the shape a promoted object takes in the body.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,059 | 10,406 | -14% | 1 | 1 | 0% | 1,627 | 2,161 | +33% | 0 | 0 | — |
case-02 | fail→pass | 12,248 | 8,069 | -34% | 1 | 1 | 0% | 1,851 | 2,034 | +10% | 0 | 0 | — |
case-03 | pass→pass | 11,796 | 19,534 | +66% | 1 | 1 | 0% | 1,663 | 1,906 | +15% | 0 | 0 | — |
case-04 | pass→pass | 18,742 | 9,235 | -51% | 1 | 1 | 0% | 2,476 | 2,186 | -12% | 0 | 0 | — |
case-05 | pass→pass | 13,578 | 10,336 | -24% | 1 | 1 | 0% | 2,212 | 2,207 | -0% | 0 | 0 | — |
case-06 | fail→fail | 24,563 | 20,639 | -16% | 1 | 1 | 0% | 3,380 | 3,358 | -1% | 0 | 0 | — |
case-07 | fail→fail | 26,508 | 23,001 | -13% | 1 | 1 | 0% | 1,766 | 2,669 | +51% | 0 | 0 | — |
case-08 | fail→pass | 8,376 | 8,481 | +1% | 1 | 1 | 0% | 882 | 1,571 | +78% | 0 | 0 | — |
case-09 | fail→pass | 14,952 | 10,940 | -27% | 1 | 1 | 0% | 2,208 | 2,522 | +14% | 0 | 0 | — |
case-10 | fail→pass | 18,158 | 9,012 | -50% | 1 | 1 | 0% | 2,273 | 2,152 | -5% | 0 | 0 | — |
case-11 | pass→pass | 14,331 | 9,528 | -34% | 1 | 1 | 0% | 1,835 | 2,043 | +11% | 0 | 0 | — |
case-12 | pass→pass | 33,869 | 12,174 | -64% | 1 | 1 | 0% | 2,323 | 2,600 | +12% | 0 | 0 | — |
case-13 | fail→fail | 10,539 | 7,069 | -33% | 1 | 1 | 0% | 1,443 | 1,876 | +30% | 0 | 0 | — |
case-14 | pass→pass | 13,369 | 8,953 | -33% | 1 | 1 | 0% | 1,889 | 1,963 | +4% | 0 | 0 | — |
case-15 | pass→pass | 13,475 | 8,219 | -39% | 1 | 1 | 0% | 2,199 | 1,965 | -11% | 0 | 0 | — |
case-16 | pass→fail | 15,269 | 30,353 | +99% | 1 | 1 | 0% | 1,966 | 2,198 | +12% | 0 | 0 | — |
case-17 | fail→fail | 11,643 | 15,271 | +31% | 1 | 1 | 0% | 1,698 | 1,827 | +8% | 0 | 0 | — |
case-18 | pass→pass | 17,648 | 21,016 | +19% | 1 | 1 | 0% | 2,622 | 2,086 | -20% | 0 | 0 | — |
case-19 | fail→pass | 20,529 | 9,589 | -53% | 1 | 1 | 0% | 1,757 | 1,937 | +10% | 0 | 0 | — |
case-20 | pass→pass | 63,657 | 8,266 | -87% | 1 | 1 | 0% | 1,949 | 2,024 | +4% | 0 | 0 | — |
case-21 | pass→pass | 13,923 | 8,152 | -41% | 1 | 1 | 0% | 1,862 | 2,086 | +12% | 0 | 0 | — |
case-22 | pass→pass | 11,848 | 7,150 | -40% | 1 | 1 | 0% | 1,588 | 1,739 | +10% | 0 | 0 | — |
case-23 | pass→pass | 13,307 | 10,644 | -20% | 1 | 1 | 0% | 1,975 | 2,417 | +22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.