Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run authorized, evidence-preserving security reviews and prepare remediation inputs.
.claude/skills/griddynamics-security/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 26% | 0% |
<security>
<purpose>
Guide coding agents through safe, contextual security review. Orchestration remains external.
</purpose>
<audience>
Coding agents reviewing software, infrastructure, platforms, interfaces, hosts, or AI systems.
</audience>
<prerequisites>
</prerequisites>
<inputs>
All invocation inputs are optional: authorized scope, workspace context, available tools, and run policy. Discover only limited metadata first. Recommend missing decisions and obtain approval at the gate that needs them.
</inputs>
<secret_gate>
assets/security-secret-scan.sh against the approved roots.READ SKILL FILE assets/security-secrets.md for secret families and handling.
</secret_gate>
<authorization>
Recommend enterprise-safe targets, environment, exclusions, coverage, limits, stop conditions, tools, credentials, data flows, and active-test bounds. Explain tradeoffs. The user approves or amends material decisions.
</authorization>
<overall_flow>
After separate lifecycle remediation, require a new clean deterministic run.
</overall_flow>
<tool_contract>
Before relying on a tool, verify and record invocation, version, supported targets, license category, local/network/SaaS behavior, credential needs, data flow, and verification date. Materially unverifiable, unavailable, GUI, hosted, or bot tools are recommendation-only.
</tool_contract>
<finding_integrity>
</finding_integrity>
<outputs>
With storage approval, write sanitized artifacts under docs/security/<run-id>/:
report.mdfindings.jsonrun.jsontasks/INDEX.mdtasks/<task-id>.mdGroup tasks by remediation area plus shared root cause/fix strategy, never by location. One task file is one concise, one-shot input for a later user-invoked coding session. Never invoke, coordinate, monitor, or validate remediation.
Without storage approval, return sanitized results without committing artifacts. Keep raw scanner output under docs/security/<run-id>/raw/; never commit it. Ask the user to review and commit; never commit or delete on their behalf.
</outputs>
<templates>
READ SKILL FILE templates/security-report.md, templates/security-run.json, templates/security-finding.json, templates/security-evidence-envelope.json, templates/security-threat-model.md, templates/security-task-index.md, templates/security-remediation-task.md.
</templates>
<asset_routing>
assets/security-architecture.mdassets/security-code.md, assets/security-packages.mdassets/security-iac.md, assets/security-containers.md, assets/security-kubernetes.md, assets/security-cloud.mdassets/security-api.md, assets/security-web-dast.md, assets/security-gateways.mdassets/security-dns-recon.md, assets/security-network-pentest.md, assets/security-exfiltration.mdassets/security-host-compliance.mdassets/security-llm-ai.mdassets/security-recommend-gui-bot.md</asset_routing>
<validation_checklist>
</validation_checklist>
<pitfalls>
</pitfalls>
</security>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,818 | 79,966 | +349% | 1 | 1 | 0% | 3,126 | 2,994 | -4% | 0 | 0 | — |
case-02 | fail→fail | 40,848 | 14,665 | -64% | 1 | 1 | 0% | 8,258 | 3,119 | -62% | 0 | 0 | — |
case-11 | pass→pass | 10,794 | 5,248 | -51% | 1 | 1 | 0% | 1,706 | 2,346 | +38% | 0 | 0 | — |
case-03 | fail→fail | 15,171 | 16,348 | +8% | 1 | 1 | 0% | 1,978 | 2,730 | +38% | 0 | 0 | — |
case-04 | fail→pass | 9,321 | 15,701 | +68% | 1 | 1 | 0% | 1,771 | 4,006 | +126% | 0 | 0 | — |
case-05 | pass→pass | 12,642 | 16,663 | +32% | 1 | 1 | 0% | 2,129 | 4,394 | +106% | 0 | 0 | — |
case-06 | pass→pass | 7,623 | 7,994 | +5% | 1 | 1 | 0% | 788 | 1,989 | +152% | 0 | 0 | — |
case-12 | pass→pass | 13,967 | 10,402 | -26% | 1 | 1 | 0% | 2,002 | 3,152 | +57% | 0 | 0 | — |
case-07 | pass→pass | 11,311 | 18,242 | +61% | 1 | 1 | 0% | 1,764 | 2,864 | +62% | 0 | 0 | — |
case-08 | fail→pass | 15,026 | 17,045 | +13% | 1 | 1 | 0% | 2,260 | 3,027 | +34% | 0 | 0 | — |
case-09 | fail→fail | 10,528 | 11,015 | +5% | 1 | 1 | 0% | 1,584 | 3,320 | +110% | 0 | 0 | — |
case-10 | pass→pass | 13,088 | 8,325 | -36% | 1 | 1 | 0% | 1,983 | 2,810 | +42% | 0 | 0 | — |
case-13 | fail→pass | 16,042 | 18,968 | +18% | 1 | 1 | 0% | 2,600 | 3,833 | +47% | 0 | 0 | — |
case-14 | fail→pass | 24,353 | 4,701 | -81% | 1 | 1 | 0% | 1,655 | 1,932 | +17% | 0 | 0 | — |
case-15 | fail→pass | 14,289 | 7,088 | -50% | 1 | 1 | 0% | 2,014 | 2,535 | +26% | 0 | 0 | — |
case-16 | fail→pass | 17,631 | 12,672 | -28% | 1 | 1 | 0% | 2,965 | 3,716 | +25% | 0 | 0 | — |
case-17 | fail→pass | 13,861 | 46,188 | +233% | 1 | 1 | 0% | 2,131 | 2,525 | +18% | 0 | 0 | — |
case-18 | fail→pass | 12,919 | 4,589 | -64% | 1 | 1 | 0% | 2,105 | 2,168 | +3% | 0 | 0 | — |
case-19 | fail→pass | 8,536 | 5,534 | -35% | 1 | 1 | 0% | 1,305 | 2,347 | +80% | 0 | 0 | — |
case-20 | fail→fail | 84,652 | 56,295 | -33% | 1 | 1 | 0% | 1,254 | 2,051 | +64% | 0 | 0 | — |
case-21 | fail→pass | 18,904 | 7,488 | -60% | 1 | 1 | 0% | 1,730 | 2,413 | +39% | 0 | 0 | — |
case-22 | fail→pass | 28,660 | 14,021 | -51% | 1 | 1 | 0% | 1,669 | 2,067 | +24% | 0 | 0 | — |
case-23 | fail→pass | 16,387 | 28,862 | +76% | 1 | 1 | 0% | 2,672 | 3,282 | +23% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +52 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/21/2026 | +32% |
Other measured skills in the registry, with their headline benchmark lift.