Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Applies Code Complete's scientific debugging method: STABILIZE → LOCATE → HYPOTHESIZE → EXPERIMENT → FIX → TEST → SEARCH. For active bug investigation, not QA process design or test coverage planning (use cc-quality-practices).
.claude/skills/ryanthedev-cc-debugging/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 30% | 0% |
| case-15 | ✓→✗ | ▼ Worse | -4% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 69% | 0% |
Most debugging time goes to finding and understanding the defect; the fix is usually obvious once you understand it. Every action tests a hypothesis — no guessing, no random changes.
First action on any bug: run the failing test or repro and read its actual output. Do this before reading the source in depth and before editing anything — the observed failure, not the code you infer it from, is what you debug. If the usual runner is unavailable (e.g. no pytest), fall back to a runnable form (python3 -c, a direct script) and capture that output now, not after a fix.
Shared numeric thresholds (20:1 debugger variation, ~50% of fixes wrong first time): Read(${CLAUDE_PLUGIN_ROOT}/references/cc-foundations.md).
STABILIZE → LOCATE → HYPOTHESIZE → EXPERIMENT → FIX → TEST → SEARCHGet a reliable reproduction — you cannot debug what you cannot reproduce.
Precondition on editing: the first tool action of the session is running the failing test or repro, and its output is captured, BEFORE any Edit to implementation code. The order is run-test → then edit, never edit-then-test. If the standard runner is missing, run the repro another way (python3 -c, a direct script) and capture that — the missing runner does not excuse skipping the observed failure.
Narrow the suspicious region before forming a hypothesis.
Form a specific, testable hypothesis — not "the bug is somewhere in module X."
Design a test that will disprove the hypothesis, not confirm it.
Fix the root cause, not the symptom.
Verify the fix actually works.
Defects cluster — if this bug existed, similar ones likely exist nearby.
Precondition on completing: before reporting the fix complete, a search for the same defect pattern (grep/Glob) has been run and its result recorded.
Rule these out before deep investigation:
< vs <=), array index vs length== instead of epsilon comparison)When quick checks fail and the systematic method stalls, the brute-force techniques (full code review, isolate in a harness, rewrite the section) and the full defect catalog are in Read(${CLAUDE_SKILL_DIR}/checklists.md).
Explain the problem out loud, to a person or a rubber duck. Articulation frequently reveals the bug before the listener responds.
| After | Next | |---|---| | Root cause found | Fix + add regression test (Steps 6–7 above) | | Defect is in untested legacy code | Skill(code-foundations:welc-legacy-code) — get it under test first | | Fix requires structural refactoring | Skill(code-foundations:cc-refactoring-guidance) |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,306 | 4,983 | -21% | 1 | 1 | 0% | 215 | 1,316 | +512% | 0 | 0 | — |
case-02 | fail→fail | 25,296 | 7,390 | -71% | 1 | 1 | 0% | 3,922 | 1,344 | -66% | 0 | 0 | — |
case-03 | fail→fail | 15,356 | 9,308 | -39% | 1 | 1 | 0% | 2,938 | 1,357 | -54% | 0 | 0 | — |
case-04 | fail→fail | 4,947 | 7,833 | +58% | 1 | 1 | 0% | 163 | 1,480 | +808% | 0 | 0 | — |
case-05 | fail→pass | 12,047 | 10,160 | -16% | 1 | 1 | 0% | 1,854 | 2,506 | +35% | 0 | 0 | — |
case-06 | fail→fail | 8,573 | 9,357 | +9% | 1 | 1 | 0% | 1,321 | 2,531 | +92% | 0 | 0 | — |
case-07 | fail→fail | 9,378 | 9,622 | +3% | 1 | 1 | 0% | 1,466 | 1,894 | +29% | 0 | 0 | — |
case-08 | pass→fail | 7,109 | 5,835 | -18% | 1 | 1 | 0% | 1,145 | 1,486 | +30% | 0 | 0 | — |
case-09 | pass→pass | 11,228 | 3,450 | -69% | 1 | 1 | 0% | 1,838 | 1,655 | -10% | 0 | 0 | — |
case-10 | fail→fail | 9,914 | 7,858 | -21% | 1 | 1 | 0% | 1,410 | 2,371 | +68% | 0 | 0 | — |
case-11 | fail→fail | 9,920 | 10,149 | +2% | 1 | 1 | 0% | 1,548 | 2,644 | +71% | 0 | 0 | — |
case-12 | pass→pass | 5,045 | 2,939 | -42% | 1 | 1 | 0% | 798 | 1,548 | +94% | 0 | 0 | — |
case-13 | pass→pass | 7,922 | 3,483 | -56% | 1 | 1 | 0% | 1,279 | 1,561 | +22% | 0 | 0 | — |
case-14 | fail→pass | 8,478 | 3,614 | -57% | 1 | 1 | 0% | 1,438 | 1,573 | +9% | 0 | 0 | — |
case-15 | pass→fail | 8,314 | 7,151 | -14% | 1 | 1 | 0% | 1,487 | 1,425 | -4% | 0 | 0 | — |
case-16 | pass→pass | 14,817 | 5,887 | -60% | 1 | 1 | 0% | 2,248 | 1,945 | -13% | 0 | 0 | — |
case-17 | pass→pass | 15,738 | 11,167 | -29% | 1 | 1 | 0% | 2,626 | 2,826 | +8% | 0 | 0 | — |
case-18 | pass→pass | 19,463 | 3,471 | -82% | 1 | 1 | 0% | 1,791 | 1,562 | -13% | 0 | 0 | — |
case-19 | fail→fail | 7,271 | 5,619 | -23% | 1 | 1 | 0% | 1,108 | 2,044 | +84% | 0 | 0 | — |
case-20 | pass→pass | 13,422 | 7,460 | -44% | 1 | 1 | 0% | 1,957 | 2,220 | +13% | 0 | 0 | — |
case-21 | pass→fail | 5,232 | 6,217 | +19% | 1 | 1 | 0% | 799 | 1,352 | +69% | 0 | 0 | — |
case-22 | pass→fail | 10,104 | 6,395 | -37% | 1 | 1 | 0% | 1,716 | 1,328 | -23% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 15 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.