Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deploy a bounded, profile-loaded adversarial review squad at an SDD point-cut (post-spec, post-plan, post-tasks, pre-merge, or an ad-hoc decision) so independent doctrine lenses converge on findings one reviewer would miss. Triggers: "deploy a squad", "adversarial squad", "post-tasks anti-laziness pass", "pre-spec investigation squad", "brownfield check", "second opinion on this design", "review squad", "run a multi-lens review". Does NOT handle: the implement-review loop (use spec-kitty-impleme
.claude/skills/priivacy-ai-adversarial-squad/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -36% | 0% |
The operational HOW. The doctrinal WHEN/WHY is the procedure adversarial-squad-deployment (packs/built-in/procedures/adversarial-squad-deployment.procedure.yaml), which sits under the brownfield-onboarding paradigm. This skill changes no mission type or guard; it is a technique the orchestrator opts into.
A squad is worth its tokens at a high-leverage point-cut where one reviewer's blind spot is expensive:
/spec-kitty.specify → pre-spec investigation (scope, prior art, live repros)/spec-kitty.plan → post-planning brownfield check (foldable issues, split-brain, deprecations)/spec-kitty.tasks → post-tasks anti-laziness pass (fakeable DoDs, decomposition realism)Do NOT use it as a rubber stamp, and do NOT wire it as a mandatory gate.
not a vibe check.
architect-alphonso — structure / seams / topologydebugger-debbie — live-evidence, coverage, "would this catch the regression?"reviewer-renata — anti-laziness, contract-vs-implementation, fakeable assertionsrandy-reducer — duplication / dead code (⚠ duct-tape bias — read critically)paula-patterns — decomposition, boundaries, second-opinion adjudicationplanner-priti — scope, sequencing, tracker hygienepython-pedro — implementer feasibilitydoctrine-daphne — doctrine integrity / DRG wiringScale past 4 only for an explicit "audit / comprehensive" ask.
"FIRST run spec-kitty agent profile show <id> and spec-kitty charter context --action <action> --json; apply the resolved initialization, boundaries, directives, and tactics, then state which you applied." Loading the profile — not naming a persona — is the point. Only a read-only harness that cannot invoke the CLI may read packs/built-in/agent_profiles/<id>.agent.yaml; that degraded fallback can diverge because overlays, specializes_from lineage, and enhances/overrides semantics are not applied. Keep delegates read-only unless the task is an isolated implementation in its own worktree.
[SEVERITY] file:line — issue — recommendation, ending in a verdict, grounded in cited evidence, with honest concession of where its lens does not apply. A steelman that over-claims is weak; an adversary that concedes nothing is noise.
lighter tier for mechanical/tracker delegates.
consequential point, do NOT average — adjudicate from the source, or dispatch one focused second-opinion delegate. Be critical of any delegate with a known bias. If irreconcilable, escalate to the human with both positions.
findings doc, or memory). The value is convergent evidence that survived independent scrutiny — not a single opinion.
Invoke by name (adversarial-squad) with the point-cut + question, e.g. "adversarial-squad: post-tasks anti-laziness on WP01–WP08." This skill is the alias surface; the doctrine procedure is the canonical record of the technique.
Bounded (3–4) · profile-LOADED · structured output · model discipline · live-evidence-grounded · second-opinion on divergence · never a mission gate.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,347 | 15,354 | -37% | 1 | 1 | 0% | 4,430 | 3,781 | -15% | 0 | 0 | — |
case-02 | fail→fail | 23,157 | 6,675 | -71% | 1 | 1 | 0% | 4,013 | 1,512 | -62% | 0 | 0 | — |
case-03 | fail→pass | 20,398 | 14,896 | -27% | 1 | 1 | 0% | 3,455 | 3,643 | +5% | 0 | 0 | — |
case-04 | fail→pass | 17,029 | 8,756 | -49% | 1 | 1 | 0% | 2,894 | 2,642 | -9% | 0 | 0 | — |
case-05 | pass→pass | 15,020 | 7,292 | -51% | 1 | 1 | 0% | 2,588 | 2,450 | -5% | 0 | 0 | — |
case-06 | pass→pass | 9,740 | 4,807 | -51% | 1 | 1 | 0% | 1,493 | 1,799 | +20% | 0 | 0 | — |
case-07 | fail→pass | 15,824 | 1,908 | -88% | 1 | 1 | 0% | 1,522 | 1,355 | -11% | 0 | 0 | — |
case-08 | fail→pass | 13,135 | 7,395 | -44% | 1 | 1 | 0% | 2,384 | 2,388 | +0% | 0 | 0 | — |
case-09 | pass→pass | 8,189 | 5,453 | -33% | 1 | 1 | 0% | 1,450 | 1,972 | +36% | 0 | 0 | — |
case-10 | pass→pass | 6,881 | 3,195 | -54% | 1 | 1 | 0% | 1,160 | 1,541 | +33% | 0 | 0 | — |
case-11 | pass→pass | 10,111 | 3,705 | -63% | 1 | 1 | 0% | 1,680 | 1,745 | +4% | 0 | 0 | — |
case-12 | fail→pass | 12,335 | 2,238 | -82% | 1 | 1 | 0% | 2,218 | 1,416 | -36% | 0 | 0 | — |
case-13 | pass→pass | 14,497 | 6,469 | -55% | 1 | 1 | 0% | 2,552 | 2,194 | -14% | 0 | 0 | — |
case-14 | pass→pass | 9,750 | 3,173 | -67% | 1 | 1 | 0% | 1,629 | 1,587 | -3% | 0 | 0 | — |
case-15 | fail→pass | 8,382 | 3,458 | -59% | 1 | 1 | 0% | 1,492 | 1,568 | +5% | 0 | 0 | — |
case-16 | fail→pass | 7,637 | 2,153 | -72% | 1 | 1 | 0% | 1,319 | 1,417 | +7% | 0 | 0 | — |
case-17 | pass→pass | 9,639 | 4,737 | -51% | 1 | 1 | 0% | 1,620 | 1,914 | +18% | 0 | 0 | — |
case-18 | pass→pass | 10,729 | 4,436 | -59% | 1 | 1 | 0% | 1,786 | 1,632 | -9% | 0 | 0 | — |
case-19 | fail→pass | 14,613 | 3,291 | -77% | 1 | 1 | 0% | 2,555 | 1,680 | -34% | 0 | 0 | — |
case-20 | fail→pass | 7,633 | 2,567 | -66% | 1 | 1 | 0% | 1,246 | 1,535 | +23% | 0 | 0 | — |
case-21 | fail→pass | 9,319 | 1,992 | -79% | 1 | 1 | 0% | 1,545 | 1,447 | -6% | 0 | 0 | — |
case-22 | fail→pass | 9,034 | 2,249 | -75% | 1 | 1 | 0% | 1,458 | 1,436 | -2% | 0 | 0 | — |
case-23 | pass→pass | 11,580 | 5,417 | -53% | 1 | 1 | 0% | 1,867 | 1,929 | +3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.