Install any skill in seconds. Free to start, no credit card required.
Get Started Free →An enterprise-grade, AI-powered penetration testing automation CLI tool. Orchestrates multiple specialized AI agents (Planner, ToolAgent, Analyst, Reporter) backed by 4 AI providers (OpenAI, Claude, Gemini, OpenRouter) and 19 integrated security tools through YAML-defined workflows. Produces professional Markdown, HTML, or JSON security reports with full evidence capture and traceability.
.claude/skills/itamarzand88-guardian-cli/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 797% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 167% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 194% | 0% |
<!-- source: guardian-cli — https://raw.githubusercontent.com/zakirkun/guardian-cli/main/SKILL.md -->
Guardian (v2.0) is a Python 3.11+ CLI application that automates penetration testing workflows using a multi-agent AI system. It is designed for authorized security assessments only.
guardian-cli/
├── ai/ # AI provider integrations & prompt templates
│ ├── providers/ # base_provider, openai, claude, gemini, openrouter
│ └── prompt_templates/
├── cli/ # CLI entry-point (Typer) & commands
│ └── commands/ # init, scan, recon, analyze, report, workflow, ai, models
├── core/ # Multi-agent orchestration engine
│ ├── agent.py # BaseAgent
│ ├── planner.py # PlannerAgent – decides next test step
│ ├── tool_agent.py # ToolAgent – selects & executes tools
│ ├── analyst_agent.py # AnalystAgent – interprets tool output
│ ├── reporter_agent.py # ReporterAgent – generates final reports
│ ├── memory.py # PentestMemory, ToolExecution, Finding dataclasses
│ └── workflow.py # WorkflowEngine – top-level orchestrator
├── tools/ # 19 security-tool wrappers (one Python file each)
├── workflows/ # YAML workflow definitions (8 built-in)
├── utils/ # logger, scope_validator, helpers
├── config/ # guardian.yaml configuration file
├── reports/ # Output directory for generated reports & session state
└── docs/ # Guides (WORKFLOW_GUIDE, TOOLS_DEVELOPMENT_GUIDE, …)Target Input
│
▼
WorkflowEngine.run_workflow() ──or── WorkflowEngine.run_autonomous()
│
├──► PlannerAgent.decide_next_action() — Strategic AI reasoning
│
├──► ToolAgent.execute_tool() — Runs the chosen security tool
│
├──► AnalystAgent.interpret_output() — Parses & links findings to executions
│
└──► ReporterAgent.execute() — Generates markdown / HTML / JSON reportEach agent inherits from BaseAgent and uses a shared PentestMemory object that stores:
| Store | Class | Purpose | |---|---|---| | findings | Finding | Vulnerabilities discovered | | tool_executions | ToolExecution | Full command + raw output | | completed_actions | list[str] | Phase progress tracker | | current_phase | str | reconnaissance → scanning → analysis → reporting |
All providers implement the same BaseProvider interface, making them interchangeable at runtime:
| Provider | Env Var | Default Model | |---|---|---| | openai | OPENAI_API_KEY | gpt-4o | | claude | ANTHROPIC_API_KEY | claude-3-5-sonnet-20241022 | | gemini | GOOGLE_API_KEY | gemini-2.5-pro | | openrouter | OPENROUTER_API_KEY | anthropic/claude-3.5-sonnet |
Switch provider via config/guardian.yaml or --provider CLI flag.
Run with python -m cli.main <command> (or guardian <command> after installation).
| Command | Purpose | |---|---| | init | Create/validate config/guardian.yaml | | scan | One-shot vulnerability scan on a target | | recon | Passive / active reconnaissance | | analyze | Re-analyze an existing session | | report | Generate / re-generate a report for a session | | workflow list | List available workflows | | workflow run | Execute a named workflow against a target | | ai | Query AI about a finding or custom prompt | | models | List configured AI providers and models | | version | Show version |
bash--target <IP/domain/CIDR> # Required for scan/recon/workflow run --provider <openai|claude|gemini|openrouter> --name <workflow-name> # For workflow run --format <markdown|html|json> --session <SESSION_ID> # For report regeneration
bash# List available workflows python -m cli.main workflow list # Web penetration test python -m cli.main workflow run --name web_pentest --target https://target.example.com # Full network assessment python -m cli.main workflow run --name network --target 192.168.1.0/24 # Autonomous AI-driven pentest python -m cli.main workflow run --name autonomous --target example.com
| File | Name | Description | |---|---|---| | recon.yaml | recon | Passive + active reconnaissance | | web_pentest.yaml | web_pentest | HTTP discovery, vuln scan, report | | network_pentest.yaml | network | Port scan, service detect, analysis | | advanced_recon.yaml | advanced_recon | Deep subdomain + DNS enumeration | | full_vuln_scan.yaml | full_vuln_scan | Comprehensive vulnerability sweep | | wordpress_audit.yaml | wordpress_audit | WordPress-specific audit | | autonomous.yaml | autonomous | AI-decides-everything mode |
yamlname: my_workflow description: "Short description" steps: - name: http_discovery type: tool # tool | analysis | report tool: httpx # tool name (must match tools/ wrapper) objective: "Describe what to find" parameters: # Override config/guardian.yaml defaults tech_detect: true threads: 100 - name: analyze type: analysis agent: analyst objective: "Correlate findings" - name: generate_report type: report # format defaults to config output.format settings: max_parallel_tools: 3 require_confirmation: true save_intermediate: true
Parameter Priority: Workflow YAML > config/guardian.yaml > Tool defaults
The engine resolves --name web → web_pentest.yaml automatically using substring matching on the filename stem.
| Category | Tools | |---|---| | Network | nmap, masscan | | Web Recon | httpx, whatweb, wafw00f | | Subdomain / DNS | subfinder, amass, dnsrecon | | Vulnerability | nuclei, nikto, sqlmap, wpscan | | SSL/TLS | testssl, sslyze | | Content Discovery | gobuster, ffuf, arjun | | Security Analysis | xsstrike, gitleaks, cmseek |
Each tool has a self-contained Python wrapper in tools/<toolname>.py that:
asyncio subprocess){"success": bool, "command": str, "raw_output": str, "exit_code": int, "duration": float}Guardian works with a subset of tools available; the AI adapts based on what is installed.
config/guardian.yaml)Key sections:
yamlai: provider: openai # Active provider openai: model: gpt-4o api_key: null # Or set OPENAI_API_KEY env var temperature: 0.2 max_tokens: 8000 pentest: safe_mode: true # Disable destructive actions require_confirmation: true max_parallel_tools: 3 tool_timeout: 300 # seconds output: format: markdown # markdown | html | json save_path: ./reports verbosity: normal # quiet | normal | verbose | debug scope: blacklist: # Never scan these - 127.0.0.0/8 - 10.0.0.0/8 - 172.16.0.0/12 - 192.168.0.0/16
Every scan session produces:
| File | Contents | |---|---| | reports/report_<SESSION_ID>.md | Full penetration test report | | reports/session_<SESSION_ID>.json | Raw session state (findings, executions, phase) |
Evidence capture includes:
execution_id field on Finding)include_reasoning: true)Report formats are selected via the output.format config key or the --format CLI flag and map to file extensions .md, .html, .json.
tools/mytool.py inheriting from BaseTool (see tools/base_tool.py)async def run(self, target: str, **kwargs) -> dict; return the standard result dicttools/__init__.pytool: mytool)See docs/TOOLS_DEVELOPMENT_GUIDE.md for full documentation.
ai/providers/myprovider_provider.py inheriting BaseProvider (ai/providers/base_provider.py)async def complete(self, messages, system_prompt) -> dictai/ai_client.py provider factoryconfig/guardian.yaml under ai:bash# Setup python -m venv venv .\venv\Scripts\activate # Windows pip install -e ".[dev]" # Run python -m cli.main --help # Test pytest tests/ # Lint / Format black . ruff check .
Core dependencies: typer[all], rich, langchain, langchain-google-genai, langchain-openai, langchain-anthropic, pyyaml, python-dotenv, pydantic, asyncio, aiofiles, jinja2
> ⚠️ Guardian is designed exclusively for authorized security testing and educational purposes. > You are fully responsible for obtaining explicit written permission before testing any system. > Unauthorized access is illegal (CFAA, GDPR, and equivalent laws worldwide).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,368 | 12,209 | +66% | 1 | 1 | 0% | 536 | 4,810 | +797% | 0 | 0 | — |
case-02 | fail→pass | 7,059 | 7,378 | +5% | 1 | 1 | 0% | 1,731 | 4,628 | +167% | 0 | 0 | — |
case-03 | fail→pass | 25,267 | 5,496 | -78% | 1 | 1 | 0% | 1,537 | 4,082 | +166% | 0 | 0 | — |
case-04 | fail→pass | 8,523 | 3,779 | -56% | 1 | 1 | 0% | 1,798 | 3,567 | +98% | 0 | 0 | — |
case-05 | fail→pass | 17,398 | 2,119 | -88% | 1 | 1 | 0% | 1,066 | 3,129 | +194% | 0 | 0 | — |
case-06 | pass→pass | 5,508 | 1,338 | -76% | 1 | 1 | 0% | 1,429 | 2,988 | +109% | 0 | 0 | — |
case-07 | pass→pass | 3,140 | 2,289 | -27% | 1 | 1 | 0% | 555 | 3,112 | +461% | 0 | 0 | — |
case-08 | fail→pass | 9,333 | 2,593 | -72% | 1 | 1 | 0% | 1,927 | 3,209 | +67% | 0 | 0 | — |
case-09 | fail→pass | 6,400 | 2,549 | -60% | 1 | 1 | 0% | 1,390 | 3,163 | +128% | 0 | 0 | — |
case-10 | pass→pass | 5,742 | 3,512 | -39% | 1 | 1 | 0% | 1,265 | 3,454 | +173% | 0 | 0 | — |
case-11 | fail→pass | 6,897 | 3,689 | -47% | 1 | 1 | 0% | 1,581 | 3,096 | +96% | 0 | 0 | — |
case-12 | fail→pass | 7,953 | 2,820 | -65% | 1 | 1 | 0% | 1,559 | 3,316 | +113% | 0 | 0 | — |
case-13 | pass→pass | 3,783 | 2,820 | -25% | 1 | 1 | 0% | 781 | 2,871 | +268% | 0 | 0 | — |
case-14 | fail→pass | 9,413 | 1,745 | -81% | 1 | 1 | 0% | 1,785 | 2,985 | +67% | 0 | 0 | — |
case-15 | fail→pass | 5,691 | 1,313 | -77% | 1 | 1 | 0% | 1,086 | 2,924 | +169% | 0 | 0 | — |
case-16 | fail→pass | 10,078 | 2,981 | -70% | 1 | 1 | 0% | 1,949 | 3,344 | +72% | 0 | 0 | — |
case-17 | fail→pass | 8,077 | 2,808 | -65% | 1 | 1 | 0% | 1,483 | 3,220 | +117% | 0 | 0 | — |
case-18 | fail→pass | 9,536 | 1,780 | -81% | 1 | 1 | 0% | 1,669 | 3,002 | +80% | 0 | 0 | — |
case-19 | fail→pass | 13,262 | 3,834 | -71% | 1 | 1 | 0% | 2,396 | 3,375 | +41% | 0 | 0 | — |
case-20 | fail→pass | 8,591 | 2,352 | -73% | 1 | 1 | 0% | 1,523 | 2,994 | +97% | 0 | 0 | — |
case-21 | pass→pass | 2,637 | 2,973 | +13% | 1 | 1 | 0% | 528 | 3,263 | +518% | 0 | 0 | — |
case-22 | pass→pass | 15,920 | 18,821 | +18% | 1 | 1 | 0% | 1,704 | 4,403 | +158% | 0 | 0 | — |
case-23 | pass→pass | 12,093 | 9,877 | -18% | 1 | 1 | 0% | 431 | 3,102 | +620% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +70 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.