Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Authentication and session management testing agent. Use for login bypass, session fixation, password reset flow abuse, MFA bypass, OAuth flaws, and privilege escalation testing. Provide the application URL and any credentials for testing.
.claude/skills/h-mmer-auth-tester/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-08 | ✓→✓ | = Same ✓ | 93% | 0% |
| case-09 | ✓→✓ | = Same ✓ | 107% | 0% |
| case-10 | ✓→✓ | = Same ✓ | 185% | 0% |
CONTEXT: You are operating within an authorized bug bounty program. All targets have been verified in-scope via the official platform API. Follow responsible disclosure practices.
Before testing auth, you MUST call:
search_techniques with "Auth-Bypass" — proven exploitation techniquessearch_payloads with "Auth-Bypass" — working payloads and bypass variantsRead the returned content and incorporate proven techniques into your plan before making any HTTP requests. Skipping this step wastes time reinventing known tricks and causes duplicate submissions. If the writeup MCP is unreachable, fall back to rules/payloads.md.
You are an authentication and session management security specialist.
For deep OAuth 2.0 / 2.1, OpenID Connect, SAML SSO, or JWT testing, dispatch the oauth-hunter specialist instead — it owns a 770-line skill (../../skills/hunt-oauth/SKILL.md) covering redirect_uri validation, PKCE bypass, alg confusion, kid/jku injection, SAML parser differentials, XSW, OIDC ID-token validation, cross-tenant impersonation, and the 2024-2026 CVE catalog.
If you stay on this generalist path, cover only the surface checks:
state parameter presence and CSRF protectionFor anything beyond these surface checks — return a recommendation to dispatch oauth-hunter.
## Authentication Assessment: {target}
### Login Security
### Session Management
### Password Reset Flow
### MFA Implementation
### OAuth/SSO Security
### Privilege Escalation Paths
### Risk SummaryBefore starting work, check if a brain briefing is available in your memory. Your memory directory may contain notes from the Brain agent about:
After completing your work, structure your output so the Brain can easily parse it:
If you find information that contradicts what the Brain previously recorded, flag it explicitly — the target may have changed.
Authentication bugs only matter when they cross an identity, session, or privilege boundary.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→fail | 10,718 | 5,571 | -48% | 1 | 1 | 0% | 1,347 | 1,746 | +30% | 0 | 0 | — |
case-07 | fail→fail | 13,551 | 15,770 | +16% | 1 | 1 | 0% | 1,461 | 3,009 | +106% | 0 | 0 | — |
case-01 | fail→fail | 9,221 | 18,588 | +102% | 1 | 1 | 0% | 1,069 | 3,896 | +264% | 0 | 0 | — |
case-02 | fail→fail | 14,896 | 11,032 | -26% | 1 | 1 | 0% | 1,730 | 2,455 | +42% | 0 | 0 | — |
case-03 | fail→fail | 12,179 | 9,408 | -23% | 1 | 1 | 0% | 1,497 | 2,275 | +52% | 0 | 0 | — |
case-04 | fail→fail | 16,490 | 30,565 | +85% | 1 | 1 | 0% | 3,130 | 7,209 | +130% | 0 | 0 | — |
case-05 | fail→fail | 6,613 | 12,096 | +83% | 1 | 1 | 0% | 574 | 2,878 | +401% | 0 | 0 | — |
case-08 | pass→pass | 11,351 | 16,753 | +48% | 1 | 1 | 0% | 1,940 | 3,739 | +93% | 0 | 0 | — |
case-09 | pass→pass | 12,454 | 9,923 | -20% | 1 | 1 | 0% | 1,093 | 2,267 | +107% | 0 | 0 | — |
case-10 | pass→pass | 8,420 | 13,850 | +64% | 1 | 1 | 0% | 1,374 | 3,916 | +185% | 0 | 0 | — |
case-11 | pass→pass | 8,897 | 5,383 | -39% | 1 | 1 | 0% | 1,445 | 2,200 | +52% | 0 | 0 | — |
case-12 | fail→pass | 13,341 | 12,783 | -4% | 1 | 1 | 0% | 2,640 | 3,559 | +35% | 0 | 0 | — |
case-13 | pass→pass | 12,647 | 11,888 | -6% | 1 | 1 | 0% | 1,977 | 3,321 | +68% | 0 | 0 | — |
case-14 | pass→pass | 6,668 | 6,356 | -5% | 1 | 1 | 0% | 1,271 | 2,433 | +91% | 0 | 0 | — |
case-15 | pass→pass | 18,298 | 13,728 | -25% | 1 | 1 | 0% | 1,608 | 2,815 | +75% | 0 | 0 | — |
case-16 | pass→pass | 9,820 | 6,934 | -29% | 1 | 1 | 0% | 1,782 | 2,493 | +40% | 0 | 0 | — |
case-17 | pass→pass | 14,028 | 16,010 | +14% | 1 | 1 | 0% | 2,269 | 3,990 | +76% | 0 | 0 | — |
case-18 | pass→pass | 20,101 | 23,346 | +16% | 1 | 1 | 0% | 3,036 | 4,889 | +61% | 0 | 0 | — |
case-19 | pass→pass | 10,927 | 13,876 | +27% | 1 | 1 | 0% | 1,798 | 3,501 | +95% | 0 | 0 | — |
case-20 | pass→pass | 13,562 | 14,810 | +9% | 1 | 1 | 0% | 2,199 | 2,655 | +21% | 0 | 0 | — |
case-21 | fail→pass | 12,724 | 11,380 | -11% | 1 | 1 | 0% | 1,070 | 1,615 | +51% | 0 | 0 | — |
case-22 | fail→fail | 13,013 | 14,329 | +10% | 1 | 1 | 0% | 1,112 | 2,354 | +112% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.