Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design and again at writing when the source reports a grid — variants crossed with backbones, datasets or metrics — and you are about to fill part of it. Covers reproducing the whole grid at reduced N where you must, and why a labelled reduced-N cell beats an empty one.
.claude/skills/tangxiangru-information-fill-the-whole-results-grid/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 401% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 20% | 0% |
The contribution of an AI/ML systems paper is a grid: method variants x backbones or base models x datasets x metric families, plus one ablation per named component and the qualitative demonstrations. Completeness of the grid is audited; nothing pays for extra depth in one cell. At design time, write the grid out as a table of empty cells and schedule the cheapest run that fills each one.
Widening is additive, not a substitution: the item the task actually ships keeps its own named subsection with its own values, in the source's units, even when the full grid gives a tighter interval. A run that priced the two scopes honestly, chose the fifteen-paper corpus, and left the one shipped paper as an appendix row scored 5/15/5 where an agent that simply printed the shipped paper's own result scored 32/25/45. See the-supplied-item-is-the-graded-unit.
Treat a single supplied example file as a smoke-test fixture for the grid, not as the evaluation set and not as something the report may drop: obtain the released implementation, the pretrained weights, the full benchmark suite and the baseline systems from the public release. When a cell cannot be run at full scale, run it at reduced N or one seed and label it as such. A crude arm counts; a Limitations sentence declaring the arm out of scope reads as the experiment never having been attempted.
Ablate each named component one at a time and report its metric delta; report every algorithmic variant of the same component side by side; use the source's own names and abbreviations throughout.
Report the sub-field's full metric set per cell rather than one scalar: a threshold metric and a ranking metric for classification, a clustering-quality metric alongside accuracy for representation claims, the paired before/after for interventions. Then decompose each aggregate per class, per difficulty stratum and per regime (in-distribution, unseen, low-data), with mean and standard deviation over seeds, and include an off-the-shelf general-purpose model as a baseline with its latency and parameter count beside its accuracy.
Information's failure is breadth: the two heaviest criteria of one task were the same experiment on two backbones and the agent ran one arm well and put the other in Limitations (46 and 0). It also fixes the 'half a metric pair' loss (F1 given, clustering metric omitted -> 12) and the 'scoped to the one supplied exemplar' loss, where every task ships one file against a full benchmark grid. Converting 'out of scope' into a labelled reduced-N arm turns structural zeros into partial credit, which is where the 43% absent rate lives.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 32,907 | 66,283 | +101% | 1 | 1 | 0% | 5,681 | 8,923 | +57% | 0 | 0 | — |
case-02 | fail→pass | 82,509 | 27,228 | -67% | 1 | 1 | 0% | 1,146 | 5,740 | +401% | 0 | 0 | — |
case-03 | fail→fail | 180,629 | 71,453 | -60% | 1 | 1 | 0% | 8,263 | 8,862 | +7% | 0 | 0 | — |
case-04 | pass→pass | 15,528 | 48,007 | +209% | 1 | 1 | 0% | 2,260 | 3,677 | +63% | 0 | 0 | — |
case-05 | fail→pass | 19,359 | 12,068 | -38% | 1 | 1 | 0% | 2,468 | 2,276 | -8% | 0 | 0 | — |
case-06 | fail→pass | 15,570 | 15,686 | +1% | 1 | 1 | 0% | 2,358 | 2,796 | +19% | 0 | 0 | — |
case-07 | pass→pass | 39,246 | 32,985 | -16% | 1 | 1 | 0% | 2,452 | 3,599 | +47% | 0 | 0 | — |
case-08 | fail→pass | 12,727 | 11,296 | -11% | 1 | 1 | 0% | 1,922 | 2,314 | +20% | 0 | 0 | — |
case-09 | fail→pass | 18,102 | 14,092 | -22% | 1 | 1 | 0% | 2,677 | 3,149 | +18% | 0 | 0 | — |
case-10 | pass→pass | 17,326 | 16,087 | -7% | 1 | 1 | 0% | 2,714 | 3,314 | +22% | 0 | 0 | — |
case-11 | pass→pass | 40,318 | 7,855 | -81% | 1 | 1 | 0% | 2,836 | 1,613 | -43% | 0 | 0 | — |
case-12 | fail→pass | 19,051 | 22,303 | +17% | 1 | 1 | 0% | 2,768 | 4,084 | +48% | 0 | 0 | — |
case-13 | fail→pass | 15,593 | 13,939 | -11% | 1 | 1 | 0% | 2,353 | 2,668 | +13% | 0 | 0 | — |
case-14 | fail→fail | 13,900 | 15,618 | +12% | 1 | 1 | 0% | 2,088 | 2,850 | +36% | 0 | 0 | — |
case-15 | fail→fail | 18,597 | 14,251 | -23% | 1 | 1 | 0% | 2,461 | 2,419 | -2% | 0 | 0 | — |
case-16 | pass→pass | 19,096 | 50,713 | +166% | 1 | 1 | 0% | 2,586 | 2,968 | +15% | 0 | 0 | — |
case-17 | pass→pass | 15,817 | 24,543 | +55% | 1 | 1 | 0% | 2,138 | 3,019 | +41% | 0 | 0 | — |
case-18 | pass→pass | 62,500 | 14,166 | -77% | 1 | 1 | 0% | 2,147 | 2,584 | +20% | 0 | 0 | — |
case-19 | fail→pass | 30,108 | 13,661 | -55% | 1 | 1 | 0% | 2,262 | 2,567 | +13% | 0 | 0 | — |
case-20 | pass→pass | 32,977 | 17,790 | -46% | 1 | 1 | 0% | 2,912 | 3,258 | +12% | 0 | 0 | — |
case-21 | pass→pass | 45,637 | 12,640 | -72% | 1 | 1 | 0% | 2,299 | 2,586 | +12% | 0 | 0 | — |
case-22 | pass→pass | 13,746 | 39,838 | +190% | 1 | 1 | 0% | 2,041 | 7,908 | +287% | 0 | 0 | — |
case-23 | pass→fail | 18,332 | 43,084 | +135% | 1 | 1 | 0% | 2,373 | 8,776 | +270% | 0 | 0 | — |
case-24 | pass→pass | 22,044 | 28,695 | +30% | 1 | 1 | 0% | 3,167 | 5,059 | +60% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 23 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +33 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.