Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user explicitly asks to fix and verify a validated or plausible security finding. Do not use as the primary trigger for full PR, commit, branch, patch, or repository scans.
.claude/skills/cowork-os-fix-finding/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 214% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 157% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 155% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -4% | 0% |
Turn a security finding into a minimal, validated code change when the issue still exists. If the issue is already fixed, prove that with focused validation and report that no code change was needed. The result should include the fix if one was needed, focused regression tests or another repeatable validation check, proof that normal behavior still works, and proof that the original issue no longer reproduces.
Start by extracting any available finding details:
If a critical field is missing, inspect the repository to fill it from code evidence. Ask the user only when the fix would otherwise require guessing a product policy or security invariant.
Use this guidance whenever reproducing the finding, running tests, or validating the fix:
AGENTS.md, README.md, setup docs, test docs, build files, or package-manager metadata for the necessary requirements.Use this checklist before calling the fix complete:
In the final response, include:
If using a scan artifact directory, resolve it using ../../references/scan-artifacts.md, then write a visible report to the fix report path. If there is no existing scan directory, a final chat summary is sufficient unless the user asks for a file.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,209 | 26,144 | +44% | 1 | 1 | 0% | 2,545 | 5,268 | +107% | 0 | 0 | — |
case-02 | fail→pass | 14,445 | 28,969 | +101% | 1 | 1 | 0% | 1,782 | 5,595 | +214% | 0 | 0 | — |
case-15 | fail→fail | 4,979 | 4,496 | -10% | 1 | 1 | 0% | 192 | 1,653 | +761% | 0 | 0 | — |
case-03 | fail→pass | 19,091 | 32,623 | +71% | 1 | 1 | 0% | 2,607 | 6,687 | +157% | 0 | 0 | — |
case-04 | fail→fail | 6,496 | 8,158 | +26% | 1 | 1 | 0% | 641 | 2,058 | +221% | 0 | 0 | — |
case-05 | fail→fail | 24,786 | 5,499 | -78% | 1 | 1 | 0% | 3,726 | 1,725 | -54% | 0 | 0 | — |
case-06 | fail→fail | 19,412 | 34,787 | +79% | 1 | 1 | 0% | 3,036 | 6,206 | +104% | 0 | 0 | — |
case-07 | fail→fail | 4,925 | 3,756 | -24% | 1 | 1 | 0% | 271 | 1,677 | +519% | 0 | 0 | — |
case-08 | fail→fail | 4,939 | 7,250 | +47% | 1 | 1 | 0% | 251 | 1,867 | +644% | 0 | 0 | — |
case-09 | pass→fail | 17,906 | 20,050 | +12% | 1 | 1 | 0% | 1,942 | 1,864 | -4% | 0 | 0 | — |
case-10 | pass→pass | 13,746 | 22,702 | +65% | 1 | 1 | 0% | 1,373 | 3,982 | +190% | 0 | 0 | — |
case-11 | fail→fail | 9,207 | 5,287 | -43% | 1 | 1 | 0% | 1,417 | 1,852 | +31% | 0 | 0 | — |
case-12 | fail→fail | 8,390 | 58,939 | +602% | 1 | 1 | 0% | 310 | 5,087 | +1541% | 0 | 0 | — |
case-13 | fail→fail | 16,529 | 4,614 | -72% | 1 | 1 | 0% | 3,119 | 1,718 | -45% | 0 | 0 | — |
case-14 | fail→fail | 5,857 | 14,849 | +154% | 1 | 1 | 0% | 158 | 1,702 | +977% | 0 | 0 | — |
case-16 | pass→fail | 7,849 | 4,680 | -40% | 1 | 1 | 0% | 1,519 | 1,800 | +18% | 0 | 0 | — |
case-17 | fail→pass | 21,712 | 24,210 | +12% | 1 | 1 | 0% | 1,734 | 4,415 | +155% | 0 | 0 | — |
case-18 | fail→fail | 4,252 | 6,256 | +47% | 1 | 1 | 0% | 143 | 1,689 | +1081% | 0 | 0 | — |
case-19 | pass→fail | 8,122 | 9,682 | +19% | 1 | 1 | 0% | 1,492 | 1,685 | +13% | 0 | 0 | — |
case-20 | pass→fail | 12,352 | 4,339 | -65% | 1 | 1 | 0% | 2,136 | 1,740 | -19% | 0 | 0 | — |
case-21 | pass→fail | 15,795 | 4,133 | -74% | 1 | 1 | 0% | 2,866 | 1,662 | -42% | 0 | 0 | — |
case-22 | fail→fail | 22,261 | 25,561 | +15% | 1 | 1 | 0% | 1,924 | 4,797 | +149% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 9 counted toward the lift figure. The other 13 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 9 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.