Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Security audit workflow - vulnerability scan → verification
.claude/skills/parcadei-security/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 6 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 143% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -54% | 0% |
Dedicated security analysis for sensitive code.
┌─────────┐ ┌───────────┐
│ aegis │───▶│ arbiter │
│ │ │ │
└─────────┘ └───────────┘
Security Verify
audit fixes| # | Agent | Role | Output | |---|-------|------|--------| | 1 | aegis | Comprehensive security scan | Vulnerability report | | 2 | arbiter | Verify fixes, run security tests | Verification report |
The /review workflow focuses on code quality. Security needs:
Task(
subagent_type="aegis",
prompt="""
Security audit: [SCOPE]
Scan for:
**Injection Attacks:**
- SQL injection
- Command injection
- XSS (Cross-Site Scripting)
- LDAP injection
**Authentication/Authorization:**
- Broken authentication
- Session management issues
- Privilege escalation
- Insecure direct object references
**Data Protection:**
- Sensitive data exposure
- Hardcoded secrets/credentials
- Insecure cryptography
- Missing encryption
**Configuration:**
- Security misconfigurations
- Default credentials
- Verbose error messages
- Missing security headers
**Dependencies:**
- Known vulnerable packages
- Outdated dependencies
- Supply chain risks
Output: Detailed report with:
- Severity (CRITICAL/HIGH/MEDIUM/LOW)
- Location (file:line)
- Description
- Remediation steps
"""
)Task(
subagent_type="arbiter",
prompt="""
Verify security fixes: [SCOPE]
Run:
- Security-focused tests
- Dependency audit (npm audit, pip audit)
- Re-check reported vulnerabilities
- Verify fixes don't introduce regressions
Output: Verification report
"""
)User: /security
→ Scan entire codebaseUser: /security authentication
→ Focus on auth-related codeUser: /security src/api/auth.py
→ Deep dive on one fileUser: /security --deps
→ Only dependency vulnerabilitiesUser: /security the payment processing code
Claude: Starting /security audit for payment code...
Phase 1: Security audit...
[Spawns aegis]
┌─────────────────────────────────────────────────────────────┐
│ Security Audit Report │
├─────────────────────────────────────────────────────────────┤
│ Scope: src/services/payment/ │
│ Files scanned: 12 │
│ Lines analyzed: 2,847 │
├─────────────────────────────────────────────────────────────┤
│ CRITICAL (1) │
│ ────────── │
│ [C1] SQL Injection in payment.py:89 │
│ query = f"SELECT * FROM orders WHERE id = {order_id}" │
│ Fix: Use parameterized queries │
│ │
│ HIGH (2) │
│ ──────── │
│ [H1] Hardcoded API key in stripe_client.py:12 │
│ STRIPE_KEY = "sk_live_..." │
│ Fix: Move to environment variable │
│ │
│ [H2] Missing input validation in refund.py:45 │
│ amount = request.json['amount'] # No validation │
│ Fix: Validate amount is positive number │
│ │
│ MEDIUM (1) │
│ ────────── │
│ [M1] Verbose error messages in error_handler.py:23 │
│ return {"error": str(e)} # Leaks internal details │
│ Fix: Return generic message, log details internally │
│ │
│ LOW (0) │
│ ─────── │
│ None │
├─────────────────────────────────────────────────────────────┤
│ Summary: 1 critical, 2 high, 1 medium, 0 low │
│ Status: BLOCKING - Fix critical issues before release │
└─────────────────────────────────────────────────────────────┘
Fix the critical and high issues, then run:
/security --verifyUser: /security --verify
Claude: Phase 2: Verifying fixes...
[Spawns arbiter]
✅ C1: SQL injection fixed - using parameterized queries
✅ H1: API key moved to environment variable
✅ H2: Input validation added
✅ M1: Error messages sanitized
All security tests passing.
Security audit: PASSED| Risk | Checked | |------|---------| | A01 Broken Access Control | ✅ | | A02 Cryptographic Failures | ✅ | | A03 Injection | ✅ | | A04 Insecure Design | ✅ | | A05 Security Misconfiguration | ✅ | | A06 Vulnerable Components | ✅ | | A07 Auth Failures | ✅ | | A08 Data Integrity Failures | ✅ | | A09 Logging Failures | ✅ | | A10 SSRF | ✅ |
--deps: Dependencies only--verify: Re-run after fixes--owasp: Explicit OWASP Top 10 report--secrets: Focus on secret detection| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,016 | 7,407 | -8% | 1 | 1 | 0% | 992 | 2,325 | +134% | 0 | 0 | — |
case-02 | fail→fail | 6,051 | 2,382 | -61% | 1 | 1 | 0% | 1,331 | 1,780 | +34% | 0 | 0 | — |
case-03 | fail→fail | 6,534 | 6,168 | -6% | 1 | 1 | 0% | 826 | 2,081 | +152% | 0 | 0 | — |
case-04 | fail→fail | 2,358 | 2,271 | -4% | 1 | 1 | 0% | 465 | 1,758 | +278% | 0 | 0 | — |
case-05 | pass→pass | 7,092 | 14,510 | +105% | 1 | 1 | 0% | 1,184 | 3,636 | +207% | 0 | 0 | — |
case-06 | pass→fail | 16,388 | 5,816 | -65% | 1 | 1 | 0% | 3,719 | 1,711 | -54% | 0 | 0 | — |
case-07 | fail→fail | 13,927 | 6,483 | -53% | 1 | 1 | 0% | 1,839 | 2,302 | +25% | 0 | 0 | — |
case-08 | fail→fail | 9,724 | 6,025 | -38% | 1 | 1 | 0% | 800 | 2,007 | +151% | 0 | 0 | — |
case-09 | fail→fail | 17,331 | 18,067 | +4% | 1 | 1 | 0% | 2,312 | 3,789 | +64% | 0 | 0 | — |
case-10 | fail→pass | 11,044 | 8,052 | -27% | 1 | 1 | 0% | 1,864 | 2,442 | +31% | 0 | 0 | — |
case-11 | fail→fail | 12,698 | 14,993 | +18% | 1 | 1 | 0% | 961 | 2,592 | +170% | 0 | 0 | — |
case-12 | fail→pass | 19,339 | 14,016 | -28% | 1 | 1 | 0% | 3,151 | 3,900 | +24% | 0 | 0 | — |
case-13 | fail→fail | 6,934 | 8,146 | +17% | 1 | 1 | 0% | 604 | 2,031 | +236% | 0 | 0 | — |
case-14 | fail→fail | 8,017 | 11,047 | +38% | 1 | 1 | 0% | 1,393 | 2,068 | +48% | 0 | 0 | — |
case-15 | fail→fail | 9,363 | 6,001 | -36% | 1 | 1 | 0% | 1,015 | 2,135 | +110% | 0 | 0 | — |
case-16 | fail→fail | 5,592 | 4,398 | -21% | 1 | 1 | 0% | 998 | 1,831 | +83% | 0 | 0 | — |
case-17 | fail→pass | 15,479 | 13,930 | -10% | 1 | 1 | 0% | 868 | 2,111 | +143% | 0 | 0 | — |
case-18 | fail→fail | 7,796 | 5,348 | -31% | 1 | 1 | 0% | 1,052 | 2,002 | +90% | 0 | 0 | — |
case-19 | fail→fail | 7,588 | 5,224 | -31% | 1 | 1 | 0% | 846 | 2,041 | +141% | 0 | 0 | — |
case-20 | pass→pass | 7,784 | 2,548 | -67% | 1 | 1 | 0% | 1,447 | 1,808 | +25% | 0 | 0 | — |
case-21 | fail→fail | 17,536 | 14,578 | -17% | 1 | 1 | 0% | 1,772 | 2,773 | +56% | 0 | 0 | — |
case-22 | fail→pass | 4,982 | 3,292 | -34% | 1 | 1 | 0% | 954 | 1,937 | +103% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.