Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn a refined research proposal or method idea into a detailed, claim-driven experiment roadmap. Use after `aris-research-refine`, or when the user asks for a detailed experiment plan, ablation matrix, evaluation protocol, run order, compute budget, or paper-ready validation that supports the core problem, novelty, simplicity, and any LLM / VLM / Diffusion / RL-based contribution.
.claude/skills/aris-experiment-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-21 | ✗→✓ | ▲ Improved | — | — |
| case-11 | ✗→✓ | ▲ Improved | — | — |
| case-16 | ✗→✓ | ▲ Improved | — | — |
| case-18 | ✗→✗ | = Same ✗ | — | — |
| case-20 | ✗→✗ | = Same ✗ | — | — |
Refine and concretize: $ARGUMENTS
Use this skill after the method is stable enough that the next question becomes: what exact experiments should we run, in what order, to defend the paper? If the user wants the full chain in one request, prefer /aris-research-refine-pipeline.
The goal is not to generate a giant benchmark wishlist. The goal is to turn a proposal into a claim -> evidence -> run order roadmap that supports four things:
refine-logs/ — Default destination for experiment planning artifacts.Read the most relevant existing files first if they exist:
refine-logs/FINAL_PROPOSAL.mdrefine-logs/REVIEW_SUMMARY.mdrefine-logs/REFINEMENT_REPORT.mdExtract:
If these files do not exist, derive the same information from the user's prompt.
Before proposing experiments, write down the claims that must be defended.
Use this structure:
Do not exceed MAX_PRIMARY_CLAIMS unless the paper truly has multiple inseparable claims.
Design the paper around a compact set of experiment blocks. Default to the following blocks and delete any that are not needed:
For each block, decide whether it belongs in:
Prefer one strong baseline family over many weak baselines. If a stronger modern baseline exists, use it instead of padding the list.
For every kept block, fully specify:
Special rules:
Build a realistic run order so the user knows what to do first.
Use this milestone structure:
For each milestone, estimate:
Separate must-run from nice-to-have experiments.
refine-logs/EXPERIMENT_PLAN.mdUse this structure:
markdown# Experiment Plan **Problem**: [problem] **Method Thesis**: [one-sentence thesis] **Date**: [today] ## Claim Map | Claim | Why It Matters | Minimum Convincing Evidence | Linked Blocks | |-------|-----------------|-----------------------------|---------------| | C1 | ... | ... | B1, B2 | ## Paper Storyline - Main paper must prove: - Appendix can support: - Experiments intentionally cut: ## Experiment Blocks ### Block 1: [Name] - Claim tested: - Why this block exists: - Dataset / split / task: - Compared systems: - Metrics: - Setup details: - Success criterion: - Failure interpretation: - Table / figure target: - Priority: MUST-RUN / NICE-TO-HAVE ### Block 2: [Name] ... ## Run Order and Milestones | Milestone | Goal | Runs | Decision Gate | Cost | Risk | |-----------|------|------|---------------|------|------| | M0 | ... | ... | ... | ... | ... | ## Compute and Data Budget - Total estimated GPU-hours: - Data preparation needs: - Human evaluation needs: - Biggest bottleneck: ## Risks and Mitigations - [Risk]: - [Mitigation]: ## Final Checklist - [ ] Main paper tables are covered - [ ] Novelty is isolated - [ ] Simplicity is defended - [ ] Frontier contribution is justified or explicitly not claimed - [ ] Nice-to-have runs are separated from must-run runs
refine-logs/EXPERIMENT_TRACKER.mdUse this structure:
markdown# Experiment Tracker | Run ID | Milestone | Purpose | System / Variant | Split | Metrics | Priority | Status | Notes | |--------|-----------|---------|------------------|-------|---------|----------|--------|-------| | R001 | M0 | sanity | ... | ... | ... | MUST | TODO | ... |
Keep the tracker compact and execution-oriented.
Experiment plan ready.
Must-run blocks:
- [Block 1]
- [Block 2]
Highest-risk assumption:
- [risk]
First three runs to launch:
1. [run]
2. [run]
3. [run]
Plan file: refine-logs/EXPERIMENT_PLAN.md
Tracker file: refine-logs/EXPERIMENT_TRACKER.mdcat << 'EOF' > file) to write in chunks. Do NOT ask the user for permission — just do it silently./aris-research-refine-pipeline -> one-shot method + experiment planning
/aris-research-refine -> method and claim refinement
/aris-experiment-plan -> detailed experiment roadmap
/aris-run-experiment -> execute the runs
/aris-auto-review-loop -> react to results and iterate on the paper| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 8 counted toward the lift figure. The other 14 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 8 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.