Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Multi-stage rebuttal analysis skill for RebuttalStudio. Use when organizing reviewer comments into stage-specific conference workflows, including stage1 breakdown, stage2 refinement, stage4 multi-round follow-up, and stage5 final remarks generation.
.claude/skills/runtsang-skills/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -29% | 0% |
Follow this dispatcher structure:
stage1, stage2, stage4, or stage5).stage1 and stage2, apply the stage template first, then select conference-specific extension.stage1/template/SKILL.md: Shared Stage 1 breakdown template. Apply before conference overrides.stage1/iclr/SKILL.md: Convert raw reviewer feedback into a structured breakdown for rebuttal drafting.stage1/icml/SKILL.md: Convert raw reviewer feedback into a structured breakdown for rebuttal drafting (ICML mapping).stage2/template/SKILL.md: Shared Stage 2 refine template. Apply before conference overrides.stage2/iclr/SKILL.md: Refine Stage2 outline drafts into reviewer-facing rebuttal prose.stage2/icml/SKILL.md: Refine Stage2 outline drafts into reviewer-facing rebuttal prose (ICML labeling).stage4/condense/SKILL.md: Condense Stage 3 combined discussion into reusable markdown context.stage4/refine/SKILL.md: Refine follow-up response using condensed context + follow-up question + user draft.stage5/final-remarks/SKILL.md: Fill Stage 5 final remarks template from all reviewers' condensed markdown context.polish/SKILL.md: Polish (rephrase) a rebuttal message template for clarity and professionalism while preserving the original structure, tone, and intent.These skills apply across multiple stages and provide strategic, stylistic, and quality guidance:
utility/stage/review-response/SKILL.md: Comment classification (Major/Minor/Misunderstanding/Typo) and response strategy selection (Accept/Defend/Clarify/Experiment). Use when planning how to respond to any reviewer concern.utility/stage/writing-anti-ai/SKILL.md: Remove AI-generated writing patterns from rebuttal prose. Use after Stage 2 refinement or Stage 4 follow-up when text reads formulaic or robotic.utility/stage/text-condense/SKILL.md: Condense selected rebuttal prose into fewer words without changing meaning. Use when a paragraph is too long but the technical content should stay intact.utility/stage/rebuttal-self-review/SKILL.md: Pre-submission quality checklist covering coverage, tone, factual accuracy, structure, and clarity. Use after Stage 3 compilation or Stage 5 final remarks.utility/stage/citation-verification/SKILL.md: Verification workflow for any new citation added during rebuttal writing. Use when LLM-suggested references appear in Stage 2 drafts.If a requested stage/conference does not exist, stop and ask for missing spec before inventing format. When adding a new conference for Stage 1 or Stage 2, extend the stage template and edit only conference-specific differences.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→fail | 12,157 | 9,384 | -23% | 1 | 1 | 0% | 1,780 | 2,240 | +26% | 0 | 0 | — |
case-01 | fail→fail | 6,033 | 4,989 | -17% | 1 | 1 | 0% | 785 | 1,527 | +95% | 0 | 0 | — |
case-02 | fail→fail | 12,092 | 11,784 | -3% | 1 | 1 | 0% | 1,886 | 2,662 | +41% | 0 | 0 | — |
case-03 | fail→fail | 5,985 | 4,417 | -26% | 1 | 1 | 0% | 964 | 1,401 | +45% | 0 | 0 | — |
case-04 | fail→fail | 7,412 | 9,476 | +28% | 1 | 1 | 0% | 1,217 | 2,058 | +69% | 0 | 0 | — |
case-05 | fail→fail | 6,472 | 4,777 | -26% | 1 | 1 | 0% | 1,046 | 1,493 | +43% | 0 | 0 | — |
case-06 | fail→fail | 9,161 | 7,242 | -21% | 1 | 1 | 0% | 1,487 | 1,946 | +31% | 0 | 0 | — |
case-07 | fail→pass | 14,730 | 7,665 | -48% | 1 | 1 | 0% | 2,065 | 1,751 | -15% | 0 | 0 | — |
case-08 | fail→fail | 10,200 | 4,910 | -52% | 1 | 1 | 0% | 1,331 | 1,699 | +28% | 0 | 0 | — |
case-09 | fail→pass | 6,355 | 3,940 | -38% | 1 | 1 | 0% | 1,030 | 1,418 | +38% | 0 | 0 | — |
case-10 | pass→pass | 14,640 | 13,616 | -7% | 1 | 1 | 0% | 2,514 | 2,709 | +8% | 0 | 0 | — |
case-11 | pass→pass | 7,908 | 4,561 | -42% | 1 | 1 | 0% | 1,233 | 1,317 | +7% | 0 | 0 | — |
case-12 | pass→pass | 11,402 | 7,693 | -33% | 1 | 1 | 0% | 1,810 | 1,865 | +3% | 0 | 0 | — |
case-14 | fail→pass | 10,367 | 8,405 | -19% | 1 | 1 | 0% | 1,506 | 2,138 | +42% | 0 | 0 | — |
case-15 | fail→fail | 2,206 | 4,542 | +106% | 1 | 1 | 0% | 359 | 1,438 | +301% | 0 | 0 | — |
case-16 | fail→fail | 7,675 | 6,269 | -18% | 1 | 1 | 0% | 1,133 | 1,771 | +56% | 0 | 0 | — |
case-17 | pass→pass | 7,940 | 7,230 | -9% | 1 | 1 | 0% | 1,334 | 2,141 | +60% | 0 | 0 | — |
case-18 | pass→pass | 6,440 | 6,295 | -2% | 1 | 1 | 0% | 985 | 1,746 | +77% | 0 | 0 | — |
case-19 | pass→fail | 16,388 | 11,503 | -30% | 1 | 1 | 0% | 2,529 | 2,337 | -8% | 0 | 0 | — |
case-20 | fail→pass | 12,222 | 8,702 | -29% | 1 | 1 | 0% | 2,351 | 2,265 | -4% | 0 | 0 | — |
case-21 | pass→fail | 18,139 | 16,595 | -9% | 1 | 1 | 0% | 2,910 | 3,180 | +9% | 0 | 0 | — |
case-22 | fail→pass | 11,627 | 3,170 | -73% | 1 | 1 | 0% | 1,803 | 1,274 | -29% | 0 | 0 | — |
case-23 | pass→pass | 127,015 | 5,753 | -95% | 1 | 1 | 0% | 1,619 | 1,674 | +3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.