Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 122% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 209% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 278% | 0% |
Run strict, evidence-first experimental analysis for ML/AI research.
Use this skill to produce a strict analysis bundle:
analysis-report.mdstats-appendix.mdfigure-catalog.mdfigures/When the user asks for review, audit, no-write, dry-run, or when inputs are incomplete, use read-only audit mode instead of producing files or figures. In that mode, output only valid/invalid statistics, blockers, claim candidates, and what evidence is missing. If invoked by /analyze-results, the command layer may write a blocker summary, but this skill should not create figures, reports, or polished conclusions from incomplete evidence.
Do not use this skill to draft a paper Results section or a full experiment wrap-up report. Those belong to ml-paper-writing or results-report.
Results prose,pubfig / pubtab,If the user wants the complete post-experiment summary report, hand off to results-report after this bundle is ready. If the user wants publication-grade figures/tables, export parameters, publication QA, or figure/table redesign, hand off to publication-chart-skill.
If the data can be read, generate real figures. Do not stop at “recommended visualization”. Exception: in read-only audit mode, do not generate figures; describe what figure would be valid after evidence is complete.
If sample size, seeds, or raw metrics are missing, state the blocker clearly.
Do not report only best scores or only p-values.
Every major figure must have purpose, caption requirements, and post-figure interpretation notes.
This skill produces analysis artifacts; it does not write manuscript sections.
Start by identifying:
csv, json, tsv, logs),Validate:
If the comparison is not statistically valid, say so before continuing. Do not treat repeated subject × task rows, folds, windows, trials, or seeds as independent units unless the design justifies it. Common blocker: a subject × task summary table is usually a repeated-measure summary, not an independent subject-level sample. If subjects have multiple task rows or missing task cells, state that before any significance or winner claim.
Before running statistics, define the exact comparison questions:
Do not mix unrelated comparisons into one undifferentiated table.
Always produce:
mean ± std when appropriate,95% CI or another clearly justified interval,Default expectation:
See:
references/statistical-methods.mdreferences/statistical-reporting.mdProduce actual figures whenever artifacts are available.
Minimum expectation for a non-trivial analysis bundle:
Every main figure must define:
See:
references/visualization-best-practices.mdreferences/figure-interpretation.mdanalysis-report.mdSummarize:
Each claim candidate should use this shape:
md## Claim Candidates - Claim: - Source evidence: - Allowed wording: - Forbidden stronger wording: - Uncertainty: - Next check: - Decision: keep | weaken | revise | discard
stats-appendix.mdRecord:
figure-catalog.mdFor each figure, record:
Do not finish until all are true:
Results draft is included.textanalysis-output/ ├── analysis-report.md ├── stats-appendix.md ├── figure-catalog.md └── figures/ ├── figure-01-main-comparison.pdf ├── figure-02-ablation.pdf └── ...
For every major figure, answer all three questions:
If a figure cannot answer question 3, it is probably decorative rather than scientific.
Use this mode when:
Return:
Do not create analysis-output/, figures, or reports in this mode. Quarantine any statistics file whose interpretation contradicts its own p-value, test method, unit of analysis, or comparison family. Do not reuse that file for claim wording until provenance is checked.
When inputs are incomplete, say so explicitly.
Examples:
Never replace missing evidence with confident prose.
Load only what is needed:
references/statistical-methods.md - test selection and assumptionsreferences/statistical-reporting.md - minimum reporting standardreferences/visualization-best-practices.md - publication-quality figure rulesreferences/figure-interpretation.md - how to explain figures with evidencereferences/analysis-depth.md - move from observation to mechanism and decisionreferences/common-pitfalls.md - common analysis and reporting failures../research-ideation/references/research-contract.md - shared claim candidate and claim strength contractexamples/example-analysis-report.mdexamples/example-stats-appendix.mdexamples/example-figure-catalog.mdOther measured skills in the registry, with their headline benchmark lift.