Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guides safe refactoring using Code Complete's fix-first-then-refactor discipline: decides between refactor, rewrite, and fix-first; enforces separate commits; applies small-change rigor to high-error-rate one-liners. Not for diagnosing the bug itself (use cc-debugging), transforming working code's complexity (use aposd-simplifying-complexity), or untested legacy code (it routes to welc-legacy-code).
.claude/skills/ryanthedev-cc-refactoring-guidance/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 3 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 53% | 0% |
Refactoring is behavior-preserving change to working code. Three rules govern it:
| Rule | Value | Source | |---|---|---| | Small-change error rate | Peaks at 1–5 lines | Weinberg 1983 | | First-attempt success | <50% for any change | Yourdon 1986b | | Review effect on 1-line changes | 55% → 2% error rate | Freedman and Weinberg 1982 |
Shared thresholds (routine length, parameters, cohesion): Read(${CLAUDE_PLUGIN_ROOT}/references/cc-foundations.md). The full 40-item refactoring checklist: Read(${CLAUDE_SKILL_DIR}/checklists.md).
Definitions:
Refactor toward something, not just away from bad code. Find the best examples of the target pattern in this codebase, the module's conventions, and how similar refactorings were structured; match the existing pattern exactly, or — if none exists — be deliberate about the precedent you're setting. Full gate: Read(${CLAUDE_PLUGIN_ROOT}/references/pattern-reuse-gate.md).
Guide the refactoring decision and the safe process.
When production is down: fix only — no refactoring or cleanup — deploy, then refactor as a separate activity when stable. Combining the two raises complexity and error likelihood.
Refactor vs rewrite:
Safe process (priority order):
If tests fail after a refactoring: don't debug extensively — back out the change, take a smaller step, and reconsider whether you understood the code.
When prerequisites are missing:
Skill(code-foundations:welc-legacy-code) and write characterization tests first; if impossible, document expected behavior, write a manual test script, and increase review rigor.Transform code smells into clean code while preserving all observable behavior (same tests pass before and after): setup/takedown smell → encapsulated interface; tramp data → direct access or restructure; duplicated code → extracted routine; speculative "design ahead" code → remove it. Behavior changes are fixing, not refactoring — fix first.
If you mixed fix and refactor in one commit:
git reset --soft HEAD~1, separate the changes, re-commit properly.If you skipped review on a small change:
Fixing a violation post-commit still costs less than the bugs it would otherwise cause.
| After | Next | |---|---| | Refactoring complete | Skill(code-foundations:cc-control-flow-quality) | | Structure changed | Skill(code-foundations:cc-routine-and-class-design) | | Reviewing a refactoring | Skill(code-foundations:aposd-reviewing-module-design) |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,540 | 6,959 | -58% | 1 | 1 | 0% | 2,899 | 2,479 | -14% | 0 | 0 | — |
case-02 | fail→pass | 14,938 | 4,451 | -70% | 1 | 1 | 0% | 1,568 | 2,043 | +30% | 0 | 0 | — |
case-03 | pass→pass | 12,456 | 6,121 | -51% | 1 | 1 | 0% | 1,956 | 2,182 | +12% | 0 | 0 | — |
case-04 | pass→pass | 10,511 | 5,755 | -45% | 1 | 1 | 0% | 1,514 | 2,176 | +44% | 0 | 0 | — |
case-05 | pass→pass | 13,062 | 7,701 | -41% | 1 | 1 | 0% | 2,029 | 2,322 | +14% | 0 | 0 | — |
case-06 | pass→pass | 11,206 | 6,016 | -46% | 1 | 1 | 0% | 1,696 | 2,229 | +31% | 0 | 0 | — |
case-07 | pass→pass | 9,042 | 6,456 | -29% | 1 | 1 | 0% | 1,413 | 2,189 | +55% | 0 | 0 | — |
case-08 | fail→fail | 14,567 | 7,839 | -46% | 1 | 1 | 0% | 2,280 | 2,490 | +9% | 0 | 0 | — |
case-09 | pass→pass | 8,791 | 6,717 | -24% | 1 | 1 | 0% | 1,226 | 2,205 | +80% | 0 | 0 | — |
case-10 | pass→pass | 12,529 | 8,148 | -35% | 1 | 1 | 0% | 1,692 | 2,399 | +42% | 0 | 0 | — |
case-11 | pass→pass | 15,250 | 15,595 | +2% | 1 | 1 | 0% | 2,172 | 2,545 | +17% | 0 | 0 | — |
case-12 | pass→pass | 6,743 | 5,210 | -23% | 1 | 1 | 0% | 1,032 | 2,078 | +101% | 0 | 0 | — |
case-13 | pass→pass | 15,774 | 8,307 | -47% | 1 | 1 | 0% | 2,303 | 2,476 | +8% | 0 | 0 | — |
case-14 | fail→pass | 9,762 | 5,530 | -43% | 1 | 1 | 0% | 1,323 | 2,014 | +52% | 0 | 0 | — |
case-15 | fail→pass | 11,048 | 4,615 | -58% | 1 | 1 | 0% | 1,516 | 1,992 | +31% | 0 | 0 | — |
case-16 | pass→pass | 12,632 | 4,530 | -64% | 1 | 1 | 0% | 1,828 | 1,933 | +6% | 0 | 0 | — |
case-17 | pass→pass | 12,084 | 8,976 | -26% | 1 | 1 | 0% | 1,863 | 2,300 | +23% | 0 | 0 | — |
case-18 | pass→pass | 8,646 | 4,697 | -46% | 1 | 1 | 0% | 1,275 | 1,867 | +46% | 0 | 0 | — |
case-19 | pass→pass | 11,711 | 6,847 | -42% | 1 | 1 | 0% | 1,758 | 2,266 | +29% | 0 | 0 | — |
case-20 | pass→pass | 15,293 | 4,645 | -70% | 1 | 1 | 0% | 2,456 | 1,954 | -20% | 0 | 0 | — |
case-21 | pass→fail | 8,323 | 6,806 | -18% | 1 | 1 | 0% | 1,367 | 2,088 | +53% | 0 | 0 | — |
case-22 | pass→pass | 13,767 | 9,100 | -34% | 1 | 1 | 0% | 2,048 | 2,607 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.