Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used when the user asks to "analyze experimental results", "generate results section", "statistical analysis of experiments", "compare model performance", "create results visualization", or mentions connecting experimental data to paper writing. Provides comprehensive guidance for analyzing ML/AI experimental results and generating paper-ready content.
.claude/skills/inno-experiment-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-22 | ✗→✓ | ▲ Improved | — | — |
| case-01 | ✗→✓ | ▲ Improved | — | — |
| case-21 | ✓→✓ | = Same ✓ | — | — |
| case-07 | ✓→✓ | = Same ✓ | — | — |
A systematic experimental results analysis workflow connecting experimental data to paper writing.
This skill provides three core capabilities:
Use this skill when you need to:
Data Loading → Data Validation → Statistical Analysis → Visualization → Writing → Quality CheckSupported Data Formats:
Data Validation Checks:
Select appropriate tools for data loading and preliminary validation based on data format.
Basic Statistics:
Significance Tests:
Select appropriate statistical tests based on data characteristics.
Key Principles:
See references/statistical-methods.md for the complete statistical methods guide.
Comparison Dimensions:
Comparison Methods:
Systematically compare performance across different methods, ensuring fair comparison.
Publication-Quality Visualization Requirements:
Common Chart Types:
Use appropriate visualization tools to generate publication-quality figures.
See references/visualization-best-practices.md for the visualization guide.
Results Section Structure:
markdown## Results ### Overview of Main Findings [1-2 paragraphs summarizing core results] ### Experimental Setup [Brief description of experimental configuration; details in appendix] ### Performance Comparison [Comparison with baseline methods, including tables and figures] ### Ablation Study [Validate contributions of each component] ### Statistical Significance [Report statistical test results] ### Qualitative Analysis [Case studies, visualization examples]
Writing Principles:
See references/results-writing-guide.md for the complete writing guide.
Checklist:
❌ Wrong approach:
✅ Correct approach:
❌ Wrong approach:
✅ Correct approach:
❌ Wrong approach:
✅ Correct approach:
See references/common-pitfalls.md for the complete error patterns and fixes.
This skill focuses on experimental results analysis and works in tandem with the ml-paper-writing skill:
inno-experiment-analysis handles:
ml-paper-writing handles:
Workflow Integration:
Experiments complete → inno-experiment-analysis analyzes
↓
Generate analysis report and visualizations
↓
ml-paper-writing integrates into paper
↓
Complete Results sectionAfter analysis, the following are generated:
analysis-report.md)figures/)results-draft.md)Refer to the examples/ directory for complete examples:
example-analysis-report.md - Complete analysis report exampleexample-results-section.md - Paper Results section exampleThe complete analysis pipeline includes:
See the guides in the references/ directory for detailed methods and best practices.
references/statistical-methods.md - Complete statistical methods guidereferences/results-writing-guide.md - Results section writing standardsreferences/visualization-best-practices.md - Visualization best practicesreferences/common-pitfalls.md - Common errors and fixes✅ Recommended:
❌ Prohibited:
✅ Recommended:
❌ Prohibited:
✅ Recommended:
❌ Prohibited:
This skill provides a systematic experimental results analysis workflow:
Following these principles produces high-quality, reproducible experimental results analysis that meets top conference standards.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-23 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +13 percentage points is the difference between those two pass rates over the 23 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.