Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design and format publication-quality figures: chart choice, color, scales, legends, captions, reproducibility.
.claude/skills/brycewang-stanford-figures/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 64% | 0% |
This is an original Open Science Skills workflow for figure production in social-science manuscripts. It is general — apply to any figure type (line, bar, point, density, map, network, small multiples). For figures whose interpretation depends on method-specific standards, also consult the relevant sibling skill (conjoint-design, conjoint-diagnostics, list-experiment, topic-modeling, text-classification, vlm-ocr-pipeline). For end-stage QA on a finished figure set, hand off to figure-table-audit.
A good figure earns its place in the manuscript: it makes a single comparison legible, it can be read without the surrounding text, and it can be regenerated from a script. If a figure cannot do all three, it is not yet ready.
Write down, in one sentence, what the figure is supposed to let the reader see. Examples:
If you cannot state the comparison in one sentence, the figure has too many goals — split it into multiple panels or multiple figures.
Match the geometry to the comparison, not to the data type:
Avoid pie charts, 3D anything, dual-axis, and donut charts in academic figures.
0.00–1.00 or 0%–100% consistently — do not mix within a figure.rainbow / jet ramps are a publication smell; replace them.Legends are read alongside the plot, so the legend order must mirror the data's visual order — readers should never have to scan back and forth to decode a series:
The rule of thumb: the eye should be able to walk from chart to legend in the same direction it reads. Top-to-bottom on the chart maps to top-to-bottom in a vertical legend, and to left-to-right in a horizontal legend.
A reader who skims should be able to understand the figure from the caption alone. Include:
Keep abbreviations defined and units explicit.
patchwork/cowplot/gridExtra or matplotlib's subplots/gridspec — not by stitching exported PNGs in Word.When asked to design or revise a figure, produce:
# Figure Plan
Comparison: <one sentence>
Chart type: <type and why>
Geometry: <axes, scales, faceting>
Color encoding: <palette, what it encodes, accessibility check>
Legend / labeling: <direct label or legend; order matches visual order>
Caption draft: <self-contained>
Reproducibility: <script path, packages, output format and dimensions>
Open issues: <anything that needs author input — denominator choice, sample restriction, etc.>When asked to produce code, default to a single ggplot2 (R) or matplotlib + seaborn (Python) script with the theme, palette, and figure dimensions explicit at the top.
figure-table-audit once the figure set is stable.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 10,032 | 8,548 | -15% | 1 | 1 | 0% | 1,985 | 3,400 | +71% | 0 | 0 | — |
case-04 | fail→fail | 13,925 | 11,314 | -19% | 1 | 1 | 0% | 2,303 | 3,749 | +63% | 0 | 0 | — |
case-01 | fail→pass | 18,594 | 9,535 | -49% | 1 | 1 | 0% | 3,383 | 3,392 | +0% | 0 | 0 | — |
case-02 | fail→pass | 17,119 | 12,047 | -30% | 1 | 1 | 0% | 3,132 | 3,992 | +27% | 0 | 0 | — |
case-03 | fail→fail | 22,367 | 14,960 | -33% | 1 | 1 | 0% | 4,246 | 4,939 | +16% | 0 | 0 | — |
case-05 | fail→pass | 15,356 | 10,925 | -29% | 1 | 1 | 0% | 2,605 | 3,739 | +44% | 0 | 0 | — |
case-06 | pass→fail | 12,839 | 6,941 | -46% | 1 | 1 | 0% | 2,034 | 3,093 | +52% | 0 | 0 | — |
case-07 | pass→pass | 10,234 | 7,413 | -28% | 1 | 1 | 0% | 2,040 | 3,400 | +67% | 0 | 0 | — |
case-08 | pass→pass | 8,372 | 7,929 | -5% | 1 | 1 | 0% | 1,511 | 3,463 | +129% | 0 | 0 | — |
case-09 | fail→fail | 11,553 | 5,176 | -55% | 1 | 1 | 0% | 2,153 | 2,738 | +27% | 0 | 0 | — |
case-11 | pass→fail | 11,134 | 7,038 | -37% | 1 | 1 | 0% | 2,284 | 3,359 | +47% | 0 | 0 | — |
case-12 | pass→pass | 9,341 | 7,283 | -22% | 1 | 1 | 0% | 1,600 | 3,199 | +100% | 0 | 0 | — |
case-13 | pass→pass | 10,330 | 8,122 | -21% | 1 | 1 | 0% | 1,879 | 3,378 | +80% | 0 | 0 | — |
case-14 | pass→pass | 11,053 | 11,082 | +0% | 1 | 1 | 0% | 1,932 | 3,770 | +95% | 0 | 0 | — |
case-15 | pass→pass | 10,249 | 5,889 | -43% | 1 | 1 | 0% | 1,824 | 2,998 | +64% | 0 | 0 | — |
case-16 | pass→pass | 9,654 | 7,484 | -22% | 1 | 1 | 0% | 1,846 | 3,302 | +79% | 0 | 0 | — |
case-17 | pass→pass | 8,381 | 8,466 | +1% | 1 | 1 | 0% | 1,534 | 3,396 | +121% | 0 | 0 | — |
case-18 | pass→pass | 8,051 | 5,574 | -31% | 1 | 1 | 0% | 1,405 | 2,816 | +100% | 0 | 0 | — |
case-19 | fail→pass | 13,111 | 11,158 | -15% | 1 | 1 | 0% | 2,300 | 3,762 | +64% | 0 | 0 | — |
case-20 | fail→pass | 5,929 | 9,898 | +67% | 1 | 1 | 0% | 959 | 3,469 | +262% | 0 | 0 | — |
case-21 | pass→fail | 20,441 | 22,569 | +10% | 1 | 1 | 0% | 4,077 | 6,332 | +55% | 0 | 0 | — |
case-22 | pass→fail | 21,091 | 23,905 | +13% | 1 | 1 | 0% | 4,400 | 6,891 | +57% | 0 | 0 | — |
case-23 | pass→pass | 10,163 | 10,292 | +1% | 1 | 1 | 0% | 1,887 | 3,614 | +92% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.