Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when Codex is already in the attack-path-analysis phase of a security scan or the user explicitly asks to trace a security finding from source to sink and calibrate severity. Do not use as the primary trigger for full PR, commit, branch, patch, or repository scans.
.claude/skills/cowork-os-attack-path-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 141% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 146% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 186% | 0% |
Turn validated or still-plausible findings into explicit attacker stories, structured attack-path analysis facts, severity calibration, and a final reportability decision grounded in the threat model.
The path references in this skill are the default locations for this phase. If the user explicitly provides a different path for a required input or output, use the user-provided path instead of the corresponding default path referenced in this skill. If a required input is still missing, stop and ask the user for it before continuing. Use the shared scan artifact path conventions in ../../references/scan-artifacts.md.
../../references/scan-artifacts.md as the repo-specific threat-model source of truth. Start from this along with the potential findings. Both inputs are required for this workflow.reportable or survives: yes even if they were not assigned polished candidate numbers during discovery.ignore.../../references/scan-artifacts.md.../../references/scan-artifacts.md. The receipt must record the candidate id, attack-path reportability decision, attack-path facts or exact proof gap, and attack-path artifact/report reference for that candidate finding.Use this checklist before finalizing the attack-path facts or policy decision:
For the most interpretive fields, explicitly ask what repository evidence suggests the opposite and why it does or does not defeat the finding:
Look specifically for repository evidence that the path is:
Apply severity and policy calibration using references/severity-policy.md.
For each surviving finding include:
Render attack-path facts using references/attack-path-facts.md.
../../references/scan-artifacts.md, even when the final policy decision is ignore or the path remains deferred.../../references/scan-artifacts.md.-- Considerations for attack path --
-- Considerations for severity / criticality re-rating --
high and above, the impact must be materially security-relevant (for example: account takeover, auth bypass, meaningful privilege escalation, significant sensitive data exposure/exfiltration, credible RCE, or similarly severe compromise), not simply a bug or strange behaviorhigh or critical, the exploitation path and impact should be clear enough that a professional security reviewer would not need a long speculative argument to justify it.critical based on contrived, highly speculative, or edge-case-only exploit stories unless the threat model explicitly supports those conditions. Critical DEMANDS attention implying an immediate likely threat.high/critical for strange configurations or odd codebase behaviors unless there is clear evidence that an in-scope attacker can realistically exploit them for major impact.ignore (or if you have to low) for criticality purposes.ignore for criticality to mark a false-positive.Non-exhaustive examples of vulnerabilities that often support critical when evidenced in code and context:
Non-exhaustive examples of vulnerabilities that often support high when evidenced in code and context:
same-site strict cookies, auth headers, csrf tokens, PUT/PATCH/DELETE, enforced json request body content type.Strong factors that often push a plausible high up to critical:
Examples that usually should not remain high/critical without very strong proof of it leading to the class of vulnerabilities above:
High/Critical acceptance checklist (all should be true, unless the threat model strongly justifies an exception):
high/critical (not informational/low) in serious audit or bug bounty triage by a major auditing firm who puts their reputation on the line.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→fail | 13,137 | 14,367 | +9% | 1 | 1 | 0% | 1,612 | 4,888 | +203% | 0 | 0 | — |
case-01 | pass→fail | 17,241 | 10,797 | -37% | 1 | 1 | 0% | 2,295 | 4,643 | +102% | 0 | 0 | — |
case-03 | pass→pass | 7,607 | 9,247 | +22% | 1 | 1 | 0% | 1,302 | 5,064 | +289% | 0 | 0 | — |
case-04 | pass→pass | 3,971 | 5,277 | +33% | 1 | 1 | 0% | 642 | 4,326 | +574% | 0 | 0 | — |
case-05 | fail→fail | 8,055 | 5,923 | -26% | 1 | 1 | 0% | 1,217 | 4,409 | +262% | 0 | 0 | — |
case-06 | fail→pass | 12,132 | 4,243 | -65% | 1 | 1 | 0% | 1,738 | 4,188 | +141% | 0 | 0 | — |
case-07 | fail→pass | 13,998 | 3,957 | -72% | 1 | 1 | 0% | 2,063 | 4,169 | +102% | 0 | 0 | — |
case-08 | fail→pass | 15,394 | 4,777 | -69% | 1 | 1 | 0% | 2,149 | 4,226 | +97% | 0 | 0 | — |
case-09 | pass→pass | 11,628 | 5,191 | -55% | 1 | 1 | 0% | 1,672 | 4,205 | +151% | 0 | 0 | — |
case-10 | pass→pass | 7,601 | 3,739 | -51% | 1 | 1 | 0% | 1,159 | 4,134 | +257% | 0 | 0 | — |
case-11 | pass→pass | 4,777 | 3,126 | -35% | 1 | 1 | 0% | 679 | 4,079 | +501% | 0 | 0 | — |
case-12 | pass→pass | 11,968 | 8,144 | -32% | 1 | 1 | 0% | 1,756 | 4,336 | +147% | 0 | 0 | — |
case-13 | fail→fail | 13,017 | 7,419 | -43% | 1 | 1 | 0% | 1,820 | 4,559 | +150% | 0 | 0 | — |
case-14 | pass→pass | 14,008 | 7,272 | -48% | 1 | 1 | 0% | 1,941 | 4,554 | +135% | 0 | 0 | — |
case-15 | fail→pass | 11,895 | 8,639 | -27% | 1 | 1 | 0% | 1,987 | 4,890 | +146% | 0 | 0 | — |
case-16 | pass→pass | 12,129 | 5,095 | -58% | 1 | 1 | 0% | 1,747 | 4,315 | +147% | 0 | 0 | — |
case-17 | pass→pass | 13,524 | 3,500 | -74% | 1 | 1 | 0% | 1,945 | 4,067 | +109% | 0 | 0 | — |
case-18 | pass→pass | 8,740 | 3,658 | -58% | 1 | 1 | 0% | 1,280 | 4,089 | +219% | 0 | 0 | — |
case-19 | pass→pass | 16,730 | 4,879 | -71% | 1 | 1 | 0% | 2,509 | 4,169 | +66% | 0 | 0 | — |
case-20 | fail→pass | 11,110 | 6,596 | -41% | 1 | 1 | 0% | 1,674 | 4,786 | +186% | 0 | 0 | — |
case-21 | fail→fail | 6,289 | 5,380 | -14% | 1 | 1 | 0% | 1,019 | 4,400 | +332% | 0 | 0 | — |
case-22 | fail→pass | 14,203 | 3,024 | -79% | 1 | 1 | 0% | 2,241 | 4,030 | +80% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.