Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate publication-quality figures and tables from experiment results. Use when user says "画图", "作图", "generate figures", "paper figures", or needs plots for a paper.
.claude/skills/wanshuiyin-paper-figure/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 119% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 203% | 0% |
Generate all figures and tables for a paper based on: $ARGUMENTS
| Category | Can auto-generate? | Examples | |----------|-------------------|----------| | Data-driven plots | ✅ Yes | Line plots (training curves), bar charts (method comparison), scatter plots, heatmaps, box/violin plots | | Comparison tables | ✅ Yes | LaTeX tables comparing prior bounds, method features, ablation results | | Multi-panel figures | ✅ Yes | Subfigure grids combining multiple plots (e.g., 3×3 dataset × method) | | Architecture/pipeline diagrams | ❌ No — manual | Model architecture, data flow diagrams, system overviews. At best can generate a rough TikZ skeleton, but expect to draw these yourself using tools like draw.io, Figma, or TikZ | | Generated image grids | ❌ No — manual | Grids of generated samples (e.g., GAN/diffusion outputs). These come from running your model, not from this skill | | Photographs / screenshots | ❌ No — manual | Real-world images, UI screenshots, qualitative examples |
In practice: For a typical ML paper, this skill handles ~60% of figures (all data plots + tables). The remaining ~40% (hero figure, architecture diagram, qualitative results) need to be created manually and placed in figures/ before running /paper-write. The skill will detect these as "existing figures" and preserve them.
publication — Visual style preset. Options: publication (default, clean for print), poster (larger fonts), slide (bold colors)pdf — Output format. Options: pdf (vector, best for LaTeX), png (raster fallback)tab10 — Default matplotlib color cycle. Options: tab10, Set2, colorblind (deuteranopia-safe)figures/ — Output directory for generated figuresgpt-6-astra — Model used via a secondary Codex agent for figure quality review./paper-plan)figures/ or project rootIf no PAPER_PLAN.md exists, scan for data files and ask the user which figures to generate.
Parse the Figure Plan table from PAPER_PLAN.md:
markdown| ID | Type | Description | Data Source | Priority | |----|------|-------------|-------------|----------| | Fig 1 | Architecture | ... | manual | HIGH | | Fig 2 | Line plot | ... | figures/exp.json | HIGH |
Identify:
Create a shared style configuration script:
python# paper_plot_style.py — shared across all figure scripts import matplotlib.pyplot as plt import matplotlib matplotlib.rcParams.update({ 'font.size': FONT_SIZE, 'font.family': 'serif', 'font.serif': ['Times New Roman', 'Times', 'DejaVu Serif'], 'axes.labelsize': FONT_SIZE, 'axes.titlesize': FONT_SIZE + 1, 'xtick.labelsize': FONT_SIZE - 1, 'ytick.labelsize': FONT_SIZE - 1, 'legend.fontsize': FONT_SIZE - 1, 'figure.dpi': DPI, 'savefig.dpi': DPI, 'savefig.bbox': 'tight', 'savefig.pad_inches': 0.05, 'axes.grid': False, 'axes.spines.top': False, 'axes.spines.right': False, 'text.usetex': False, # set True if LaTeX is available 'mathtext.fontset': 'stix', }) # Color palette COLORS = plt.cm.tab10.colors # or Set2, or colorblind-safe def save_fig(fig, name, fmt=FORMAT): """Save figure to FIG_DIR with consistent naming.""" fig.savefig(f'{FIG_DIR}/{name}.{fmt}') print(f'Saved: {FIG_DIR}/{name}.{fmt}')
Use this decision tree for data-driven figures (inspired by Imbad0202/academic-research-skills):
| Data Pattern | Recommended Type | Size | |-------------|-----------------|------| | X=time/steps, Y=metric | Line plot | 0.48\textwidth | | Methods × 1 metric | Bar chart | 0.48\textwidth | | Methods × multiple metrics | Grouped bar / radar | 0.95\textwidth | | Two continuous variables | Scatter plot | 0.48\textwidth | | Matrix / grid values | Heatmap | 0.48\textwidth | | Distribution comparison | Box/violin plot | 0.48\textwidth | | Multi-dataset results | Multi-panel (subfigure) | 0.95\textwidth | | Prior work comparison | LaTeX table | — |
For each figure in the plan, create a standalone Python script:
Line plots (training curves, scaling):
python# gen_fig2_training_curves.py from paper_plot_style import * import json with open('figures/exp_results.json') as f: data = json.load(f) fig, ax = plt.subplots(1, 1, figsize=(5, 3.5)) ax.plot(data['steps'], data['fac_loss'], label='Factorized', color=COLORS[0]) ax.plot(data['steps'], data['crf_loss'], label='CRF-LR', color=COLORS[1]) ax.set_xlabel('Training Steps') ax.set_ylabel('Cross-Entropy Loss') ax.legend(frameon=False) save_fig(fig, 'fig2_training_curves')
Bar charts (comparison, ablation):
pythonfig, ax = plt.subplots(1, 1, figsize=(5, 3)) methods = ['Baseline', 'Method A', 'Method B', 'Ours'] values = [82.3, 85.1, 86.7, 89.2] bars = ax.bar(methods, values, color=[COLORS[i] for i in range(len(methods))]) ax.set_ylabel('Accuracy (%)') # Add value labels on bars for bar, val in zip(bars, values): ax.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 0.3, f'{val:.1f}', ha='center', va='bottom', fontsize=FONT_SIZE-1) save_fig(fig, 'fig3_comparison')
Comparison tables (LaTeX, for theory papers):
latex\begin{table}[t] \centering \caption{Comparison of estimation error bounds. $n$: sample size, $D$: ambient dim, $d$: latent dim, $K$: subspaces, $n_k$: modes.} \label{tab:bounds} \begin{tabular}{lccc} \toprule Method & Rate & Depends on $D$? & Multi-modal? \\ \midrule \citet{MinimaxOkoAS23} & $n^{-s'/D}$ & Yes (curse) & No \\ \citet{ScoreMatchingdistributionrecovery} & $n^{-2/d}$ & No & No \\ \textbf{Ours} & $\sqrt{\sum n_k d_k / n}$ & No & Yes \\ \bottomrule \end{tabular} \end{table}
Architecture/pipeline diagrams (MANUAL — outside this skill's scope):
figures/, preserve it and generate only the LaTeX \includegraphics snippet[MANUAL] in the figure plan and latex_includes.texbash# Run all figure generation scripts for script in gen_fig*.py; do python "$script" done
Verify all output files exist and are non-empty. Then render-then-verify: re-open each RENDERED PDF/PNG (not the script) and self-check — no clipped labels, no legend covering data, every number/label readable at final print size. This self-check happens BEFORE the Step 7 review, so the reviewer's budget goes to substance, not to catching clipped axes.
For each figure, output the LaTeX code to include it:
latex% === Fig 2: Training Curves === \begin{figure}[t] \centering \includegraphics[width=0.48\textwidth]{figures/fig2_training_curves.pdf} \caption{Training curves comparing factorized and CRF-LR denoising.} \label{fig:training_curves} \end{figure}
Save all snippets to figures/latex_includes.tex for easy copy-paste into the paper.
Send figure descriptions and captions to GPT-6-Astra for review:
spawn_agent:
model: gpt-6-astra
reasoning_effort: xhigh
message: |
Review these figure/table plans for a [VENUE] submission.
For each figure:
1. Is the caption informative and self-contained?
2. Does the figure type match the data being shown?
3. Is the comparison fair and clear?
4. Any missing baselines or ablations?
5. Would a different visualization be more effective?
[list all figures with captions and descriptions]The checklist is PARTITIONED (pattern from Anthropic's Claude Science figure-style skill, Apache-2.0): correctness rules always bind — they are about whether the figure tells the truth, have no aesthetic content, and no style choice may override them; guidance rules are defaults — they produce a clean result, but a deliberate, stated alternative may override them.
Correctness — always binds, verify against the DATA before the render:
source either disappears entirely or is drawn visibly distinct (open / hatched marker, named in the key); it never feeds a mean/CI plotted alongside included rows
row — if one category contradicts the claim, qualify it ("on 3 of 4 benchmarks") or downgrade to a description; a figure that overclaims is wrong even if it renders beautifully
/ protocol are not drawn as visual peers; separate them or mark the difference in the caption
says n and the unit of replication (panel or caption)
(not the script) actually happened: no clipped labels, no legend covering data, every number/label readable at final print size
Guidance — strong defaults (from pedrohcgs/claude-code-my-workflow), a deliberate stated alternative may override — EXCEPT items that Key Rules below make hard (vector-PDF output and no-titles-inside-figures are Key Rules: treat those two as binding, not overridable):
\caption{} (from pedrohcgs)emp_rate)plt.title for publications)figures/
├── paper_plot_style.py # shared style config
├── gen_fig1_architecture.py # per-figure scripts
├── gen_fig2_training_curves.py
├── gen_fig3_comparison.py
├── fig1_architecture.pdf # generated figures
├── fig2_training_curves.pdf
├── fig3_comparison.pdf
├── latex_includes.tex # LaTeX snippets for all figures
└── TABLE_*.tex # standalone table LaTeX files| Type | When to Use | Typical Size | |------|------------|--------------| | Line plot | Training curves, scaling trends | 0.48\textwidth | | Bar chart | Method comparison, ablation | 0.48\textwidth | | Grouped bar | Multi-metric comparison | 0.95\textwidth | | Scatter plot | Correlation analysis | 0.48\textwidth | | Heatmap | Attention, confusion matrix | 0.48\textwidth | | Box/violin | Distribution comparison | 0.48\textwidth | | Architecture | System overview | 0.95\textwidth | | Multi-panel | Combined results (subfigures) | 0.95\textwidth | | Comparison table | Prior bounds vs. ours (theory) | full width |
Design pattern (type × style matrix) inspired by baoyu-skills. Publication style defaults and figure rules from pedrohcgs/claude-code-my-workflow. Visualization decision tree from Imbad0202/academic-research-skills.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 41,771 | 17,338 | -58% | 1 | 1 | 0% | 7,629 | 4,278 | -44% | 0 | 0 | — |
case-02 | fail→fail | 15,358 | 13,723 | -11% | 1 | 1 | 0% | 238 | 3,892 | +1535% | 0 | 0 | — |
case-03 | fail→fail | 45,655 | 9,192 | -80% | 1 | 1 | 0% | 8,259 | 4,259 | -48% | 0 | 0 | — |
case-04 | fail→pass | 46,078 | 23,947 | -48% | 1 | 1 | 0% | 8,247 | 7,487 | -9% | 0 | 0 | — |
case-05 | fail→pass | 51,185 | 30,400 | -41% | 1 | 1 | 0% | 8,239 | 6,709 | -19% | 0 | 0 | — |
case-06 | fail→pass | 19,932 | 15,805 | -21% | 1 | 1 | 0% | 2,578 | 5,636 | +119% | 0 | 0 | — |
case-07 | fail→pass | 23,103 | 19,312 | -16% | 1 | 1 | 0% | 3,280 | 6,485 | +98% | 0 | 0 | — |
case-08 | fail→pass | 18,884 | 21,699 | +15% | 1 | 1 | 0% | 2,229 | 6,765 | +203% | 0 | 0 | — |
case-09 | fail→pass | 17,997 | 15,429 | -14% | 1 | 1 | 0% | 2,287 | 5,245 | +129% | 0 | 0 | — |
case-10 | fail→pass | 14,152 | 18,401 | +30% | 1 | 1 | 0% | 1,834 | 6,186 | +237% | 0 | 0 | — |
case-11 | fail→pass | 15,579 | 17,034 | +9% | 1 | 1 | 0% | 1,578 | 5,650 | +258% | 0 | 0 | — |
case-12 | fail→pass | 18,754 | 13,985 | -25% | 1 | 1 | 0% | 2,145 | 5,136 | +139% | 0 | 0 | — |
case-13 | pass→pass | 20,862 | 13,687 | -34% | 1 | 1 | 0% | 2,649 | 5,285 | +100% | 0 | 0 | — |
case-14 | fail→pass | 17,926 | 15,816 | -12% | 1 | 1 | 0% | 2,292 | 5,615 | +145% | 0 | 0 | — |
case-15 | fail→pass | 39,553 | 9,342 | -76% | 1 | 1 | 0% | 2,155 | 4,443 | +106% | 0 | 0 | — |
case-16 | pass→pass | 20,957 | 18,890 | -10% | 1 | 1 | 0% | 2,509 | 6,304 | +151% | 0 | 0 | — |
case-17 | pass→pass | 18,054 | 19,632 | +9% | 1 | 1 | 0% | 1,920 | 6,094 | +217% | 0 | 0 | — |
case-18 | pass→pass | 18,408 | 28,504 | +55% | 1 | 1 | 0% | 2,042 | 5,558 | +172% | 0 | 0 | — |
case-19 | pass→pass | 13,711 | 12,943 | -6% | 1 | 1 | 0% | 1,298 | 4,969 | +283% | 0 | 0 | — |
case-20 | fail→pass | 14,957 | 8,698 | -42% | 1 | 1 | 0% | 1,530 | 4,243 | +177% | 0 | 0 | — |
case-21 | fail→fail | 17,402 | 9,510 | -45% | 1 | 1 | 0% | 2,123 | 4,277 | +101% | 0 | 0 | — |
case-22 | fail→pass | 19,025 | 14,001 | -26% | 1 | 1 | 0% | 2,308 | 5,342 | +131% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/11/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.