Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Adversarial security and resilience analysis — auto-triggered during /review and /test based on task classification. Provides attack surface analysis, boundary testing, auth bypass attempts, dependency chain attacks, and Beast Mode stress testing.
.claude/skills/bilal140202-red-team-adversarial/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 33% | 0% |
This skill applies adversarial thinking to code changes: instead of checking against a compliance list (that's what security_guardrails.md does), it actively asks "how would an attacker exploit this change?" and "what breaks under extreme conditions?"
It complements — never replaces — the existing OWASP security scan in /review.
/review and /test phases only. It cannot override gates, skip phases, or alter classification.AI MUST check the task classification from the Work Log and apply this matrix automatically:
Classification │ /review │ /test
──────────────────────┼──────────────────┼─────────────────
tiny-fix │ — │ —
quick-win │ — │ —
hotfix │ Lite Red Team │ Lite Adversarial (1-2 cases)
feature │ Full Red Team │ Adversarial Cases
architecture-change │ Full Red Team │ Adversarial Cases + Beast ModeAuto-trigger logic: During /review or /test, read Classification: from the active Work Log. If classification is hotfix, feature, or architecture-change, execute the corresponding mode below. No user action required.
Minimal overhead (≤300 tokens output). Focus exclusively on the fix point:
Output: 1-2 findings max, using the Red Team Report format below.
Comprehensive adversarial analysis of all changed files:
Output: All findings using the Red Team Report format below.
In addition to Full Red Team, analyze systemic resilience:
Output: Use the Beast Mode Analysis format below.
markdown## Red Team Findings ### [CRITICAL|HIGH|MEDIUM|LOW] — [Attack Vector]: [Brief Description] - **File**: [path:line] - **Attack Scenario**: [1-2 line attack narrative] - **Impact**: [What attacker gains] - **Mitigation**: [Concrete fix]
markdown## Adversarial Test Cases | # | Category | Input / Scenario | Expected Behavior | Priority | |---|----------|------------------|--------------------|----------| | 1 | Boundary | [extreme input] | [should reject/handle] | HIGH | | 2 | AuthZ Bypass | [bypass attempt] | [should deny] | CRITICAL |
markdown## Beast Mode Analysis ### Concurrency - [race condition scenarios with file:line references] ### Resource Exhaustion - [memory/CPU/disk scenarios] ### Fault Injection - [what happens if dependency X fails? if network drops?]
For phase-entry loading, read only:
Ironclad RulesWhen to Use (Auto-Trigger Matrix)ModesLoad Output Formats, Blocking Rules, Work Log Integration, Red Team Findings, and Common Mistakes on full read or cache miss only.
Red Team findings use a graduated blocking model:
│ Security Guardrails (OWASP) │ Red Team Skill
────────────────────┼─────────────────────────────┼──────────────────
CRITICAL │ Hard block │ Hard block
HIGH │ Hard block │ Soft block (record decision)
MEDIUM │ Flag, proceed allowed │ Advisory
LOW │ Informational │ Advisory/review verdict. MUST fix before proceeding.## Red Team Findings section). Recommend using /decide to document accept/defer rationale.Rationale: Red Team analysis is inherently more speculative than OWASP checklist scanning. Hard-blocking on HIGH would create excessive false positives. But CRITICAL findings (e.g., a directly exploitable auth bypass path) must be treated as blocking.
All Red Team findings MUST be recorded in the Work Log under ## Red Team Findings:
markdown## Red Team Findings - [date] /review: [N] findings ([severity breakdown]) - [date] /test: [N] adversarial cases generated - HIGH risk decisions: [accepted/deferred — see /decide #N]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 14,276 | 9,457 | -34% | 1 | 1 | 0% | 2,548 | 3,216 | +26% | 0 | 0 | — |
case-08 | fail→pass | 11,936 | 8,508 | -29% | 1 | 1 | 0% | 2,267 | 2,920 | +29% | 0 | 0 | — |
case-09 | fail→pass | 13,682 | 5,537 | -60% | 1 | 1 | 0% | 2,138 | 2,748 | +29% | 0 | 0 | — |
case-10 | pass→pass | 13,545 | 7,553 | -44% | 1 | 1 | 0% | 2,291 | 2,797 | +22% | 0 | 0 | — |
case-11 | fail→pass | 11,238 | 3,508 | -69% | 1 | 1 | 0% | 1,932 | 2,247 | +16% | 0 | 0 | — |
case-01 | fail→fail | 17,294 | 11,367 | -34% | 1 | 1 | 0% | 2,170 | 3,015 | +39% | 0 | 0 | — |
case-03 | fail→fail | 7,964 | 16,321 | +105% | 1 | 1 | 0% | 943 | 3,274 | +247% | 0 | 0 | — |
case-04 | pass→pass | 21,145 | 17,130 | -19% | 1 | 1 | 0% | 3,803 | 4,702 | +24% | 0 | 0 | — |
case-05 | fail→pass | 12,182 | 9,336 | -23% | 1 | 1 | 0% | 2,673 | 3,565 | +33% | 0 | 0 | — |
case-06 | pass→pass | 20,368 | 17,393 | -15% | 1 | 1 | 0% | 3,172 | 4,226 | +33% | 0 | 0 | — |
case-07 | fail→pass | 16,472 | 12,693 | -23% | 1 | 1 | 0% | 2,041 | 3,020 | +48% | 0 | 0 | — |
case-12 | pass→pass | 24,549 | 12,198 | -50% | 1 | 1 | 0% | 1,088 | 2,356 | +117% | 0 | 0 | — |
case-13 | pass→pass | 10,065 | 4,315 | -57% | 1 | 1 | 0% | 1,677 | 2,369 | +41% | 0 | 0 | — |
case-14 | fail→pass | 16,375 | 1,991 | -88% | 1 | 1 | 0% | 2,996 | 1,900 | -37% | 0 | 0 | — |
case-15 | pass→pass | 15,154 | 8,875 | -41% | 1 | 1 | 0% | 2,475 | 3,034 | +23% | 0 | 0 | — |
case-16 | fail→pass | 12,994 | 3,402 | -74% | 1 | 1 | 0% | 2,050 | 2,105 | +3% | 0 | 0 | — |
case-17 | fail→pass | 14,453 | 4,329 | -70% | 1 | 1 | 0% | 2,451 | 2,074 | -15% | 0 | 0 | — |
case-18 | pass→pass | 16,304 | 3,424 | -79% | 1 | 1 | 0% | 2,451 | 2,130 | -13% | 0 | 0 | — |
case-19 | pass→pass | 19,056 | 11,323 | -41% | 1 | 1 | 0% | 3,022 | 3,266 | +8% | 0 | 0 | — |
case-20 | pass→pass | 16,896 | 3,164 | -81% | 1 | 1 | 0% | 2,335 | 1,920 | -18% | 0 | 0 | — |
case-21 | pass→pass | 21,758 | 20,614 | -5% | 1 | 1 | 0% | 2,819 | 4,285 | +52% | 0 | 0 | — |
case-22 | fail→pass | 11,824 | 6,039 | -49% | 1 | 1 | 0% | 1,958 | 2,478 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.