Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Adversarial code auditor that hunts down bugs, logic errors, and security flaws. Use for deep correctness passes, not style reviews.
.claude/skills/sickn33-bugs-are-annoying/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 138% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 123% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 78% | 0% |
An adversarial QA pass for any codebase, in any language. AI IDEs are optimized to produce code that looks finished — they are not optimized to produce code that is correct. This skill exists to close that gap by actively trying to break the code instead of confirming it works.
Treat all code as guilty until proven innocent. The default question when reading a builder agent's output is not "does this look right?" — it's "how would this break, and what did the author not think of?"
This is an adversarial pass, not a confirmatory one. Do not skim and approve. Do not skip a category because it "seems fine." Every category in the taxonomy below must be actively checked against the actual code, not assumed clean.
Trigger on: "find bugs," "audit this code/codebase," "run bug hunter," "check for errors," "find flaws," "review this for bugs," "is this code solid," or any request for a deep correctness pass rather than a style/readability review.
Do not skip phases or collapse them into a single skim. Each phase catches things the others miss.
git diff), or a specific area. Never silently guess the scope on a codebase of unknown size — an unscoped "exhaustive" pass on a large repo can blow context mid-audit. Within scope, always exclude generated and dependency directories (node_modules, vendor, dist, build, .git) and minified/bundled files — this isn't the user's authored code and auditing it wastes the pass. Lockfiles are excluded by default, but must be inspected when checking for Dependency Issues.bugs.md — Use the exact format below. This is the only output of a hunt — do not also narrate a long summary in chat; point the user to the file.Language-agnostic. Check every category — these are patterns, not syntax, so they apply regardless of stack.
Stylistic or formatting preferences are explicitly not bugs. Do not log them.
Dormant bugs: if a bug sits on a code path that isn't currently reachable or used (e.g. a variable that's computed but never read), it still gets the severity it would have if active — do not downgrade it for being unreachable. Add a one-line note to the entry that it isn't currently triggered, e.g. "Not yet triggered — finalPricePerItem is computed but unused."
bugs.mdWrite this file at the root of the project being audited (or the relevant scope if auditing a subfolder). Use this exact structure:
markdown# Bug Report — [project/scope name] — [date] ## Summary - Critical: N open, N fixed - Intermediate: N open, N fixed - Normal: N open, N fixed ## 🔴 Critical ### BUG-001: [Short title] - **File:** path/to/file.ext:line - **Issue:** what is actually wrong - **Trigger:** the exact input/sequence that causes it - **Impact:** what breaks because of it - **Suggested Fix:** described or sketched, not applied - **Confidence:** *(omit if fully confirmed in-scope; include "Needs Verification" if it depends on code outside the audited scope)* - **Status:** Open ## 🟡 Intermediate ... ## 🟢 Normal ... ## ✅ Resolved ### BUG-0XX: [Title] — Fixed [date] (kept for history, moved here once fixed)
Rules for entries:
file:line reference — never "somewhere in this file."BUG-001, BUG-002, ...), even across multiple runs.When bugs-are-annoying is run again on a codebase that already has a bugs.md:
Open bug against the current code — if it's actually fixed now, move it to ✅ Resolved with the date.The file is a running history of the codebase's health, not a disposable report.
bugs.md. Code is only changed if the user explicitly asks afterward (e.g. "fix BUG-003," "fix all Critical bugs"). Until then, every fix described in bugs.md is a suggestion only.bugs.md.Confidence: Needs Verification rather than asserting it as certain.bugs.md with the Summary counts and the date — a clean result is part of the history, not a no-op.Only enters this mode when the user explicitly asks to fix something — e.g. "fix BUG-001," "fix all Critical bugs," "apply the suggested fixes for the Intermediate ones."
bugs.md and locate the specified bug ID(s) or severity tier.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-23 | pass→pass | 8,256 | 4,779 | -42% | 1 | 1 | 0% | 1,422 | 3,223 | +127% | 0 | 0 | — |
case-01 | fail→fail | 5,917 | 12,426 | +110% | 1 | 1 | 0% | 654 | 4,364 | +567% | 0 | 0 | — |
case-02 | fail→fail | 3,414 | 3,395 | -1% | 1 | 1 | 0% | 227 | 2,609 | +1049% | 0 | 0 | — |
case-03 | fail→fail | 3,990 | 4,280 | +7% | 1 | 1 | 0% | 230 | 2,648 | +1051% | 0 | 0 | — |
case-04 | pass→pass | 10,453 | 6,696 | -36% | 1 | 1 | 0% | 2,351 | 3,884 | +65% | 0 | 0 | — |
case-05 | pass→pass | 11,849 | 13,099 | +11% | 1 | 1 | 0% | 2,664 | 5,368 | +102% | 0 | 0 | — |
case-06 | pass→fail | 13,189 | 10,701 | -19% | 1 | 1 | 0% | 2,613 | 3,788 | +45% | 0 | 0 | — |
case-07 | pass→pass | 17,699 | 14,884 | -16% | 1 | 1 | 0% | 1,350 | 3,664 | +171% | 0 | 0 | — |
case-08 | pass→pass | 7,811 | 4,922 | -37% | 1 | 1 | 0% | 1,454 | 3,287 | +126% | 0 | 0 | — |
case-09 | pass→pass | 8,692 | 3,476 | -60% | 1 | 1 | 0% | 1,580 | 2,840 | +80% | 0 | 0 | — |
case-10 | fail→pass | 6,613 | 4,223 | -36% | 1 | 1 | 0% | 1,320 | 3,146 | +138% | 0 | 0 | — |
case-11 | pass→pass | 3,090 | 2,729 | -12% | 1 | 1 | 0% | 466 | 2,847 | +511% | 0 | 0 | — |
case-12 | pass→pass | 1,648 | 1,884 | +14% | 1 | 1 | 0% | 333 | 2,754 | +727% | 0 | 0 | — |
case-13 | fail→pass | 7,286 | 2,667 | -63% | 1 | 1 | 0% | 1,465 | 2,852 | +95% | 0 | 0 | — |
case-14 | fail→pass | 8,425 | 4,786 | -43% | 1 | 1 | 0% | 1,562 | 3,476 | +123% | 0 | 0 | — |
case-15 | pass→pass | 6,805 | 3,580 | -47% | 1 | 1 | 0% | 1,212 | 3,043 | +151% | 0 | 0 | — |
case-16 | fail→pass | 6,406 | 3,314 | -48% | 1 | 1 | 0% | 1,079 | 2,944 | +173% | 0 | 0 | — |
case-17 | pass→pass | 9,133 | 4,000 | -56% | 1 | 1 | 0% | 1,511 | 2,909 | +93% | 0 | 0 | — |
case-18 | fail→fail | 8,840 | 2,019 | -77% | 1 | 1 | 0% | 1,642 | 2,746 | +67% | 0 | 0 | — |
case-19 | pass→pass | 7,443 | 4,229 | -43% | 1 | 1 | 0% | 1,398 | 3,019 | +116% | 0 | 0 | — |
case-20 | pass→pass | 9,962 | 4,598 | -54% | 1 | 1 | 0% | 1,676 | 3,120 | +86% | 0 | 0 | — |
case-21 | pass→pass | 9,352 | 15,983 | +71% | 1 | 1 | 0% | 1,568 | 3,146 | +101% | 0 | 0 | — |
case-22 | fail→pass | 10,230 | 3,856 | -62% | 1 | 1 | 0% | 1,711 | 3,050 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +17 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.