Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when turning measured results into a table or figure for the paper — building a LaTeX or markdown results table from workspace/results/*.json, deciding what uncertainty to report, choosing which baselines and ablations belong in the main table, or writing a caption that stands alone.
.claude/skills/tangxiangru-result-table/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 127% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 4% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 20% | 0% |
By Stage 06 the run has machine-readable results under workspace/results/ and an experiment_manifest.json indexing them with schema metadata. The table in the paper must be derived from those files, not retyped. A number that appears in the manuscript and nowhere in workspace/results/ is unsupported, and it is the single easiest thing for a reviewer to catch.
workspace/results/experiment_manifest.json for the result artifactsand their schemas.
workspace/code/, not by hand.A regenerable table survives a Stage 06 rerun; a hand-typed one silently goes stale the moment a result changes.
(workspace/writing/tables.tex, or an inline markdown table in the report), so the paper and the artifact cannot drift.
The main table answers the paper's one claim. Everything else goes to the appendix.
eye lands on it after the baselines.
not concern dilutes it and invites reviewers to find a column where you lose.
whose baselines are weak reads as a weak result regardless of the margin.
State what the spread means, every time. 0.74 ± 0.03 is meaningless without knowing whether that is a standard deviation, a standard error, or a confidence interval, and over how many seeds.
explicitly — see the venue-checklist skill.
present it in a format that implies replication.
within-noise win is the most common quiet overclaim in a results table.
Reviewers read figures and tables before the body. The caption must carry: what is being compared, on what data, with what metric, over how many runs, and what the reader should conclude. "Table 2: Results." fails all five.
Good: "Table 2: Accuracy on the held-out split of BENCHMARK, mean ± standard deviation over 5 seeds. Retrieval recovers 12 points that long-context prompting loses when the relevant evidence is diffuse."
The same rules apply, plus:
.png (or another renderableraster format) under report/images/ and embedded with a report-relative path — . A .pdf figure shows the report viewer nothing, and the Stage 07 gate rejects it.
workspace/figures/ and are included with apath relative to main.tex.
should be referenced. An unreferenced figure is either dead weight or a missing paragraph.
workspace/results/.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,721 | 9,755 | -9% | 1 | 1 | 0% | 459 | 1,111 | +142% | 0 | 0 | — |
case-02 | fail→fail | 11,031 | 10,685 | -3% | 1 | 1 | 0% | 435 | 1,374 | +216% | 0 | 0 | — |
case-03 | fail→fail | 9,689 | 9,620 | -1% | 1 | 1 | 0% | 261 | 1,094 | +319% | 0 | 0 | — |
case-04 | pass→pass | 10,917 | 8,342 | -24% | 1 | 1 | 0% | 1,807 | 2,161 | +20% | 0 | 0 | — |
case-05 | fail→fail | 9,598 | 2,900 | -70% | 1 | 1 | 0% | 1,710 | 1,265 | -26% | 0 | 0 | — |
case-06 | fail→fail | 17,023 | 11,468 | -33% | 1 | 1 | 0% | 1,994 | 2,067 | +4% | 0 | 0 | — |
case-07 | fail→pass | 14,791 | 11,295 | -24% | 1 | 1 | 0% | 1,728 | 1,938 | +12% | 0 | 0 | — |
case-18 | pass→pass | 9,847 | 9,641 | -2% | 1 | 1 | 0% | 1,625 | 1,493 | -8% | 0 | 0 | — |
case-17 | pass→pass | 9,819 | 9,394 | -4% | 1 | 1 | 0% | 1,379 | 1,653 | +20% | 0 | 0 | — |
case-08 | pass→pass | 9,873 | 7,016 | -29% | 1 | 1 | 0% | 1,785 | 2,169 | +22% | 0 | 0 | — |
case-09 | pass→pass | 12,105 | 11,524 | -5% | 1 | 1 | 0% | 1,215 | 2,002 | +65% | 0 | 0 | — |
case-10 | pass→pass | 10,779 | 7,190 | -33% | 1 | 1 | 0% | 1,693 | 2,021 | +19% | 0 | 0 | — |
case-11 | fail→pass | 18,453 | 6,499 | -65% | 1 | 1 | 0% | 1,807 | 2,005 | +11% | 0 | 0 | — |
case-12 | pass→pass | 11,390 | 4,668 | -59% | 1 | 1 | 0% | 1,780 | 1,639 | -8% | 0 | 0 | — |
case-13 | fail→fail | 16,886 | 12,852 | -24% | 1 | 1 | 0% | 2,008 | 2,072 | +3% | 0 | 0 | — |
case-14 | fail→pass | 19,183 | 8,179 | -57% | 1 | 1 | 0% | 991 | 2,252 | +127% | 0 | 0 | — |
case-15 | pass→pass | 9,397 | 11,176 | +19% | 1 | 1 | 0% | 1,519 | 1,922 | +27% | 0 | 0 | — |
case-16 | pass→fail | 13,721 | 13,153 | -4% | 1 | 1 | 0% | 1,854 | 1,928 | +4% | 0 | 0 | — |
case-19 | fail→fail | 16,066 | 10,694 | -33% | 1 | 1 | 0% | 1,639 | 1,780 | +9% | 0 | 0 | — |
case-20 | pass→pass | 10,576 | 13,234 | +25% | 1 | 1 | 0% | 1,436 | 2,773 | +93% | 0 | 0 | — |
case-21 | pass→pass | 20,973 | 21,278 | +1% | 1 | 1 | 0% | 4,463 | 4,291 | -4% | 0 | 0 | — |
case-22 | pass→pass | 11,255 | 13,930 | +24% | 1 | 1 | 0% | 1,683 | 2,235 | +33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.