Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at analysis and writing when a multi-stage pipeline is about to be reported by its end-to-end metric alone. Covers exhibiting each stage's intermediate object, and re-running the source's own demonstrations on the source's own inputs rather than on yours.
.claude/skills/tangxiangru-information-exhibit-the-intermediate-objects/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 132% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 553% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -13% | 0% |
In systems and modelling work the mechanism is the claim, so the intermediate objects are the evidence. For any multi-stage derivation, pipeline or multi-module framework, plan one exhibit per stage inside the report: the object that stage emits, in the source's own notation - the intermediate expression, the region proposals, the retrieved set, the routing assignment, the repaired candidate. A correct final number with the chain omitted is graded as an unshown derivation, not as a result.
Add the field's canonical diagnostics of the internal representation, not only outcome curves: a low-dimensional projection of the embeddings, a pairwise correlation or covariance heatmap, per-class feature-distribution overlap, and attention or saliency maps overlaid on the inputs - each shown before and after the intervention, on the same axes.
Reproduce qualitative demonstrations on the source's own showcase inputs and prompts, and quote the system's verbatim output: the transcript, the generated artifact, the before-to-after answer change. Your own prompt set may demonstrate the capability but does not reproduce the demonstration, and aggregate statistics never substitute for it.
Reproduce the claims in the original's framing first, then add disagreements additively - a demonstration that a claim is a protocol artifact must still carry that claim's evidence in the original's units and plot types beside it.
Everything that counts lives in the single report document: stage outputs, the derivation, and each figure's headline numbers in its caption. Work parked in a side document or a results directory, however good, is not part of the paper.
Covers the three remaining Information mechanisms. The largest single measured loss was a 162 KB derivation with all intermediate steps written to a side file while the report never used the terms at all - three criteria scored 28/18/32 for work that was done. Three criteria demand representation diagnostics (embedding projections, correlation structure, attention maps, before/after distributions) that no report drew, and two demand the showcase demonstration verbatim, where substituting the agent's own prompts scored 0. The reproduce-then-critique ordering addresses the runs that pivoted to refuting the source and earned 5-25.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 24,146 | 48,238 | +100% | 1 | 1 | 0% | 3,796 | 8,818 | +132% | 0 | 0 | — |
case-02 | fail→pass | 49,038 | 44,045 | -10% | 1 | 1 | 0% | 8,233 | 8,716 | +6% | 0 | 0 | — |
case-03 | fail→pass | 79,596 | 134,513 | +69% | 1 | 1 | 0% | 7,751 | 7,755 | +0% | 0 | 0 | — |
case-04 | fail→pass | 9,300 | 77,496 | +733% | 1 | 1 | 0% | 1,289 | 8,418 | +553% | 0 | 0 | — |
case-05 | fail→fail | 7,074 | 45,467 | +543% | 1 | 1 | 0% | 1,014 | 9,247 | +812% | 0 | 0 | — |
case-06 | pass→fail | 29,738 | 54,241 | +82% | 1 | 1 | 0% | 2,437 | 8,702 | +257% | 0 | 0 | — |
case-07 | fail→pass | 18,230 | 12,851 | -30% | 1 | 1 | 0% | 2,545 | 2,219 | -13% | 0 | 0 | — |
case-08 | fail→pass | 12,236 | 8,108 | -34% | 1 | 1 | 0% | 1,764 | 1,730 | -2% | 0 | 0 | — |
case-09 | pass→pass | 16,164 | 12,213 | -24% | 1 | 1 | 0% | 2,302 | 2,103 | -9% | 0 | 0 | — |
case-10 | fail→pass | 20,688 | 43,197 | +109% | 1 | 1 | 0% | 2,963 | 2,740 | -8% | 0 | 0 | — |
case-11 | fail→pass | 15,398 | 23,538 | +53% | 1 | 1 | 0% | 2,064 | 2,563 | +24% | 0 | 0 | — |
case-12 | pass→pass | 17,074 | 11,633 | -32% | 1 | 1 | 0% | 2,309 | 2,170 | -6% | 0 | 0 | — |
case-13 | fail→pass | 20,882 | 16,171 | -23% | 1 | 1 | 0% | 2,968 | 2,798 | -6% | 0 | 0 | — |
case-14 | fail→pass | 61,559 | 6,884 | -89% | 1 | 1 | 0% | 1,497 | 1,429 | -5% | 0 | 0 | — |
case-15 | fail→pass | 11,249 | 13,206 | +17% | 1 | 1 | 0% | 1,583 | 2,156 | +36% | 0 | 0 | — |
case-16 | fail→fail | 16,478 | 13,155 | -20% | 1 | 1 | 0% | 2,379 | 2,334 | -2% | 0 | 0 | — |
case-17 | pass→pass | 22,701 | 12,237 | -46% | 1 | 1 | 0% | 2,398 | 2,292 | -4% | 0 | 0 | — |
case-18 | fail→pass | 14,306 | 21,778 | +52% | 1 | 1 | 0% | 2,079 | 1,953 | -6% | 0 | 0 | — |
case-19 | fail→pass | 31,173 | 11,117 | -64% | 1 | 1 | 0% | 1,745 | 2,148 | +23% | 0 | 0 | — |
case-20 | fail→pass | 14,372 | 38,168 | +166% | 1 | 1 | 0% | 2,052 | 1,707 | -17% | 0 | 0 | — |
case-21 | fail→pass | 14,429 | 19,478 | +35% | 1 | 1 | 0% | 2,167 | 2,088 | -4% | 0 | 0 | — |
case-22 | fail→pass | 13,112 | 10,383 | -21% | 1 | 1 | 0% | 1,891 | 2,025 | +7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +68 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.